Dian Ang Yap
•

Post-Training NVIDIA Nemotron 3.5 Lightning on Harvey LAB
Small open models have gotten good fast. Each generation of NVIDIA Nemotron has narrowed the distance between the smallest model in the family and the largest, and in production that matters more than any other trend in open weights, because model size decides what you can afford to do: serve on every matter rather than the expensive ones, retrain nightly rather than quarterly, and maintain a separate post-trained model per firm rather than one model everyone shares.
On hard domain work, though, teams still keep reaching for the largest base they can afford. The reason is not a lack of capability in the small models. It is a cold start problem. Nemotron 3.5 Lightning at baseline passes none of our held-out Harvey Legal Agent Bench (LAB) tasks, and under all-pass grading traditionally, a model at zero is not merely behind, it is stuck: no successful trajectories to reinforce, and no signal pointing toward what the rubric wants. Post-training a model that never succeeds is a different problem from post-training one that sometimes does, and it is the problem that has kept small models out of specialized production work.
Our previous two field reports post-trained Nemotron 3 Super and Nemotron 3 Ultra on LAB and brought both into frontier range at a fraction of closed-model cost. This report takes the same harness, data, and recipe down to the smallest model in the family, where the economics are best and the cold start is hardest. Post-training ran on the Trajectory platform in a few hours, with no new engineering between one base and the next.
Nemotron 3.5 Lightning's performance on LAB
Post-trained Nemotron 3.5 Lightning reaches 8.3%, up from 0% at baseline, placing it at the top of the models we evaluated, above Opus 4.6 at 6.6% and above post-trained Nemotron 3 Ultra at 5.8%.

Baseline and post-trained all-pass rates for each Nemotron base. Nemotron 3.5 Lightning moves from 0% to 8.3%, Nemotron 3 Ultra from 0.8% to 5.8%, Nemotron 3 Super from 0.8% to 3.3%. Shaded band shows the range of the leading closed models.
One recipe lifts every base it is applied to, and the largest movement belongs to the smallest model. The highlight is that post-trained Nemotron 3.5 Lightning also lands above post-trained Nemotron 3 Ultra, which is early signs that it’s not the base intelligence that matters, but the ceiling that a model allows for. We’re pushing further on this direction by exploring how one can evaluate the post-trainability of models, are excited to share more soon.

All-pass rate on held-out LAB tasks across open and closed models, baseline and post-trained.
Where the gains land
End-to-end wins on a benchmark this coarse can concentrate in a single corner of practice, which would make them much less interesting. But here, they did not.

Change in all-pass rate by practice area after post-training, in percentage points. Nine of 24 areas improve; none regress.
Nine of 24 practice areas improved and none regressed, with gains in data privacy and cyber, arbitration and international disputes, corporate governance, emerging companies and VC, trade and sanctions, litigation, tax, trusts and estates, and white collar. At five tasks per category the resolution is coarse, so the coverage is the finding rather than the magnitude of any single bar: post-training on one firm's production work lifted the model across five unrelated families of legal task without costing it performance anywhere.
Cost is where a model this size separates from the field. Lightning runs at a fraction of the cost of Super and Ultra, and far far cheaper than most closed source alternatives.
Research preview: Intelligence Density
Right now a model is a finished object. You take it or you leave it. We think it should be shapeable, and that the people closest to the work should have the tools to shape it. Quality on your work, cost per task, latency, where the model spends its effort and where it stops: all adjustable in principle, almost none of them in practice, because moving any one takes a training stack, an evaluation loop, and enough runs to know which direction you are going. Our work is doing that research once and handing back the knob.
Intelligence density is the next one. Capability per token generated, rather than capability per parameter held. It is the property that decides what a model costs on real work, and it is the one per-token pricing misses: a model can win on price per token and lose on price per task, because nothing in the objective limits how many it spends getting there.
Measured that way, the base model does not have low density. It has none: 22k output tokens per held-out task and nothing passed. Post-training is what creates density in the first place, and the first post-trained model converts 90k tokens into ten passes.

Distribution of output tokens per held-out task. The base averages 22k, the first post-trained model 90k, and the second holds the same pass rate at 37k. Not shown: both trained models maintain the same reward.
The second run is where it compounds. We apply our intelligence density techniques and get the same ten tasks in 37k tokens, roughly 2.4x the density of the first, with no measured loss on the held-out set. All-pass grading rewards a model for checking its work and never tells it when to stop, so post-training lengthens the route to the answer even as it finds the answer more often. Shorten the path and density rises again.
A frozen model burns the same tokens on the thousandth matter as the first. A model in the loop gets denser every cycle. This is a preview rather than a full result, so stay tuned for the in-depth analysis and for how to use this yourself on the Trajectory platform.
Small models, same pipeline
Continual learning only works if the loop is cheap enough to close. Every cycle is a full pass: traces in, benchmark rebuilt, weights updated, model redeployed. Run that on a frontier model and you can afford it quarterly, for everyone at once. Run it on a small one and you can afford it nightly, per customer.
So the loop wants (and needs!) the smallest model that can do the work. Nemotron 3.5 Lightning is a promising glimpse at a future where this is possible. And because our learning layer is model-agnostic, the same loop picks up whatever ships next, at whatever size makes it cheap to keep running.
Train and customize Nemotron Lightning now on the Trajectory Platform.



