Everything Became a Training Problem
notes from open source evening @FactoryAI tonight! such a star line up of speakers and panels, learnt so much from the talks from @modal
@FactoryAI
Thanks for reading! Subscribe for free to receive new posts and support my work.
@baseten
@MiniMax_AI
@Kimi_Moonshot
The line between training and inference is basically gone.
The big three, and what training did to them
The framing from the first talk: a few years ago you could treat weights as a finished product. Download, serve, tune the config, done. Now a lot of the wins come from post training the model to be cheap to run. Faster inference gets you more data, more data gets you a better model, loop.
The big three optimizations are still quantization, KV cache, and speculation. All three now have a training step attached.
Quantization. The result that stuck with me is that quantizing more of the model can score the same or better. Pushing FP4 into all output projections at every layer, plus more than half the shared experts, held quality while adding roughly 20%. The intuition is that rounding error isn’t monotonic: error in one layer can be partially cancelled by error in the next, and calibration steers it. The analogy from the talk was a course correction. A small error at launch misses the moon by a lot. The same error in orbit still lands you on the crater.
KV cache. Compaction today is semantic. You decide what to drop and eat a cache miss when the prefix changes. The alternative is synthesis: keep a compressed representation and attend to it, produced in a single pass rather than reasoned about. The claim was over 90% compression while keeping most of the semantic information and not noticeably moving benchmarks, provided you post train the model to read it. Differentiable compressed memory the model treats as full context.
Speculation. Diffusion drafters change the shape here. You get a fixed draft length, but speculators already have a fixed draft length, so that isn’t a real tradeoff. What you gain is proposing every token in the draft at once instead of left to right, which makes drafts both faster and longer. The harder claim was that draft acceptance rate drifts with message content and sequence lengths, so you’d want to retrain the speculator continuously while serving. That needs data permission, hidden state generation, storage, all interleaved with whatever else you’re training. Only pencils out at very large deployments.
The forward look was two lines: get good at FP4 without quality loss, and treat KV cache as a hot potato. Extremely valuable, extremely big, nobody wants to hold it. So you move it within GPUs, between prefill and decode workers, over the network.
Why your coding agent should not be everyone else’sToday every customer runs the same agent on different code. A bank migrating COBOL and a startup shipping React get the same one.
Their thesis was four levers, shallow to deep: skills and memory, prompts and model selection, coordination (routing thresholds, subagent design, task decomposition), and finally weights. The argument for the deep lever is that the first three all live in the context window. Nothing compounds, and your conventions and incident history ride to a third party model call after call. Weights keep the learning scoped to your org.
The sharper version: your most valuable training data is traces, reviews, and fixes on proprietary systems, and often that data can never leave your environment. No frontier lab will ever train on it. Which means a general model is capped at general competence in your domain no matter how big it gets. The honest caveat is cost and time.
They’ve shipped the narrow version already. Small open weight models fine tuned as a guard layer for secrets and risky commands, with decision thresholds calibrated against production data using the logprobs open weights expose. Narrow task, dataset from run traces, tune, calibrate, ship.
The evals piece was my favorite part of the night. You can’t tune what you can’t measure, and public benchmarks say nothing about your specific workstream. So they mine merged changes, recover intent from the inputs and the verifier from the outputs, and assemble tasks that are exactly three things: a prompt, a verifier, and an environment.
The gating before a candidate becomes a task is where it gets good. Hermetic reproducible environment. No network, so the agent can’t go look up the merged PR. Verifier replayed from the real change rather than hand written. Covers the intent without being over or under specific. Passes on the golden patch, fails on the base commit. Then they sabotage the patch on purpose to confirm the verifier catches it, plus a flakiness check before it’s admitted.
The result is a living benchmark that evolves as the codebase does. It’s also diagnostic. One failure they showed was the model changing routing architecture but never finding the now obsolete provider logic. That’s a code search gap, fixable with a cheap lever. Not everything needs weights.
On where the human goes: today runs are human triggered and borderline cases human audited. Next, humans handle only exceptions. The endgame is scheduled self evaluation where humans set policy. Corrections get captured as signal, so reviewers become teachers rather than graders.
Multimodality that isn’t bolted onThe failure mode with encoders added after pretraining is that visual tokens get projected into the text representation and the model knows an image is there, but there’s no real causal attention between visual and text tokens. It can’t resolve which thing in the image you meant.
They tried adding visual tokens during continued pretraining and it hurt text capability. Tried introducing it earlier and hit extreme hyperparameter sensitivity at scale. Landed on native training from scratch over text and vision together: no measurable text regression, more stable training. And you can see it in the attention maps. With the post hoc encoder, the block between visual and text regions is dead. With native training, real interaction. The same approach extends to audio.
It shipped as a 30B open weight omni model. Text, image, video, audio in a shared representation, up to 12 references. A context model builds the joint representation of everything you fed it, the base model generates, then a final pass produces new 2K output rather than upscaling.
The community had it running on Apple silicon within 48 hours of release, unsupported by the lab. This keeps happening, and it’s the most underrated argument for open weights: engineering effort you could not buy.
YES,
Quantization, cache compression, speculation, agent specialization, cross modal alignment. Every one of them became a training problem.
Thanks for reading! Subscribe for free to receive new posts and support my work.