luna-swarm
A planner, a supervisor and many cheap models in one hierarchy, measured on the CommonMark spec tests.
98.3%of CommonMark spec tests passed by the hierarchy, against 52.8% for the cheap model alone
An orchestrator splits a goal into small tasks, a supervisor handles failures, and cheap models do the work. On the 652 CommonMark 0.31.2 spec tests, the cheap model alone passes 344 (52.8 percent) and the hierarchy passes 641 (98.3 percent). On the 162 tests held out, the hierarchy passes 151 (93.2 percent).
The idea under test was that task size predicts whether a cheap model succeeds. It did not hold: the planner's size labels did not predict first-try success.