Sissi Wang

why videos feel heavy?

why videos feel heavy?

yesterday i was walking in the norwegian mountains with

( update: we’re in for the tech bros cohort2 ! cooking beautiful stuff, more soon)

Thanks for reading! Subscribe for free to receive new posts and support my work.

talking about what it actually means to contribute to something. we talked about tiktok, and why that platform is able to transform how people behave over time: it plays with attention, compressing entertainment into the smallest unit our brain will accept. and then we landed on a question beneath it: what is actually stopping people from taking more videos than they do? everyone wants to capture and moment right?

videos are ‘big’ and takes up storage.

this is when compression starts to sound fun. if videos felt light, if it took up less space and moved faster, people would shoot more of them. it decides whether our videos buffer, music libraries fit in phone, and maybe, maybe, what intelligent even is.

i went down a little bit and here’s what i found, and where i think the real openings are.

what is compression, like actually actuallycompression is one idea: prediction. if you can predict what comes next, you don’t need to store it. you only need to store things that are unpredictable. A perfectly predictable signal compresses to almost nothing.

pure noise compresses to nothing, since noise is by definition, all surprise.

shannon proved that the limit of compression is the entropy of the source, and entropy is just a measure of how unpredictable something is.

a confession about encoding and decodingi thought that encoding is the easy part and decoding is the hard part. i had it exactly backwards (very interesting)

1 - encoding is the act of compression: taking a few raw things (say, a video, which is just a huge stack of photos shown fast) and rewriting it in the a much smaller form. Decoding is the reverse: taking the small form and reconstructing the pictures so your screen can show them. we shoot videos, our phone encodes it. we watch videos, our phone decodes. the pair of rules for doing both is called a codec, and names like h.264 or av1 are just generations of those rules.

2 - here is my intuition for how encoding shrinks a video. 2 neighboring frame are almost identical: only a few things moved. so instead of storing every frame in full the encoder stores one full frame and then, for the rest, only the differences: this block of pixels slid left, that region stayed the same, here is a little bit that actually changed. finidng the best way to describe those differences is a search problem, and the search is enormous. the encoder tries thousands of ways to slice up each frame and describe its motion, hunting for the cheapest description.

3 - decoding, is just following the finished recipe. hard decisions are already made and written into the file. the decode replays them. but if we think about it, it kinda needs to be easy as the economics are wildly lopsided: a movie gets encoded once, on a server farm where nobody is waiting, and then decoded millions of times on cheap phones with tiny dedicated chips.

one expensive machine, one stream of bits, and an endless wall of featherweight devices on the other side. every deployed codec in history is a version of this pic ^^

neural codecs scramble this arrangement. instead of hand-designed rules like ‘split the frame into blocks, describe each block’s motion, a neural codec uses trained neural netwroks to do the compressing the reconstructing. there are three main families, and each one distributes the work between encoder and decoder differently.

fam1, autoencoder codecs

2 networks trained together as a pair, the encoder netwrok takes a frame and squeezes it down through successive layers into a small grid of numbers called a latent, which is what actually gets stored or sent. the decoder network runs the same idea in reverse, expanding the latent back into pixels layer by layer. because the 2 networks are architectural mirror images, they cost roughly the same to run: both sides execute millions of multiply-accumulate operations per frame on a gpu or neural accelerate. compare that to the classical setup, where the decoder’s job was so simple it fit in a tiny dedicated circuit. the cheap decode that every phone on Earth relies on is gone, and playback, the side that runs a million times, now carries real compute cost.

fam2, diffusion codecs

these invert the classical asymmetry entirely. the encoder is cheap: it spits out a tiny description of the image. the decoder is a generative diffusion model (same machinery as ai image generators), which rebuilds the picture from noise, refining over many steps, each step a full network pass. so decoding costs tens of network runs per image. the payoff: at extreme compression it doesn’t reconstruct exact pixels, it generates plausible detail that looks right. that’s why they win perceptual benchmarks at bitrates where classical codecs turn into blocky mush. but for streaming, where one encode serves a million playbacks, putting the expensive side on playback is the worst possible shape.

fam3, implicit neural representations (INRs)

these flip everything. there is no general-purpose encoder at all. for each video, you train a small network from scratch to memorize that specific video: frame number in, that frame’s pixels out. the compressed file IS the network’s weights. encoding = a full training loop per video, minutes to hours of gpu. but decoding a frame is one forward pass through a deliberately tiny network. fast and simple. this recreates the good old asymmetry: pay huge once at encode, every playback stays cheap.

so the question is not ‘can we make everything easier’. count how many times each side runs. a youtube video is encoded once and decoded millions of times, so making encoding 100x harder to make decoding 2x cheaper is a fantastic trade. the reverse is a disaster. a codec is an economic contract between one encoder and many decoders, and the shape of that contract, more than any benchmark score, decides what wins.

the state of play in 2026video: av1 now powers ~30% of all netflix streaming, using one third less bandwidth than avc/hevc with 45% fewer buffering interruptions. its successor av2 dropped in june 2026, ~30% bitrate savings over av1. here’s what surprised me: av2 is NOT an ai codec. in a world drowning in neural everything, the flagship next-gen video codec is a conventional design, because it has to run on fixed-function silicon in a billion devices. meanwhile vvc/h.266, technically excellent, is stuck in patent purgatory: five years after finalization, not one smartphone ships a native hardware decoder. standards die of licensing more often than they die of math.

neural video codecs are winning benchmarks and losing deployments: microsoft’s dcvc-rt (cvpr 2025) hit 125 fps encode / 113 fps decode at 1080p, ~21% bitrate savings over the vvc reference, plus integer arithmetic so different chips decode the same bits the same way. that last part sounds boring and is actually everything. floating point differs slightly across chips, and in a codec a slightly different answer means the video turns to soup. classical codecs solved bit-exactness decades ago. neural ones are re-learning it the hard way.

images: jpeg, born 1992, still carries most of the web. jpeg ai became the first end-to-end learned image standard in 2025. and jpeg xl pulled off a full redemption arc: kicked out of chrome in 2022, championed by apple and the pdf association, re-added to chrome 145 in feb 2026 via a memory-safe rust decoder (still behind a flag). adoption is a social process wearing a technical costume.

audio: neural audio codecs (soundstream, encodec, descript’s dac) quietly found their real product-market fit, and it’s not bandwidth. their discrete tokens are the interface layer for speech and music language models. the codec became a tokenizer. nobody planned this.

the wild one: a 2025 nature machine intelligence paper, lmcompress, used large generative models as predictors feeding arithmetic coding and roughly DOUBLED the lossless ratios of jpeg-xl for images, flac for audio, h.264 for video, and quadrupled bz2 on text. remember the first principle: prediction is compression. the better a model understands data, the smaller the file. the catch is speed: multi-billion-parameter models compress at kilobytes per second. a formula 1 car you have to push by hand.

where the actual gaps are(imo)the deployment gap. neural codecs already beat classical ones on quality-per-bit across basically every modality. they’re barely deployed anyway, because:

1. no fixed-function hardware. classical codecs get dedicated decode blocks in every chip; neural codecs fight for the gpu and drain the battery.

2. cross-device determinism. floating point breaks bit-exact reconstruction.

3. installed base and standards inertia. jpeg has reigned for over thirty years.

I THINK WE SHOULD:

look hard at INRs. the nerv line of work reframes video compression as model compression: overfit a small network that maps frame index to frame, then compress the network. decoding becomes a cheap forward pass, 38 to 132x faster than pixel-wise approaches, and hinerv pushed quality into genuinely competitive territory. INRs recreate the asymmetry that made codecs deployable: spend arbitrary compute once, decode cheaply everywhere. most deployment-plausible neural approach, still immature. exactly the combination you want.

look at underserved data. video and images have armies. eeg, ecg, fmri, microscopy volumes, simulation output, time-series: the incumbents there (sz, zfp) are classical compressors reaching about 2:1 lossless on scientific floats, and learned methods are barely getting started. a domain where the right quality metric is not psnr but ‘does the decompressed signal still support the diagnosis’ is a domain where a thoughtful newcomer gets to define the rules.

right-size the lmcompress insight. the full-size version is undeployable, but the principle scales down. a small predictor, fine-tuned for one vertical (genomics, logs, medical text), feeding arithmetic coding, could capture most of the gain at usable speed. the gap between ‘doubles every known ratio’ and ‘kilobytes per second’ is not a footnote. it’s a business plan waiting for someone patient.

(what we should NOT do: build a general-purpose video or image codec. that arena belongs to google, netflix, meta, apple, qualcomm, samsung, and microsoft research. hardware, patents, distribution: a wall, not a hill.)

the thought i could not shakethe hutter prize has paid out for years on a simple bet: compressing wikipedia well is equivalent to understanding it. deepmind showed a language model trained only on text compresses images and audio better than png and flac. lmcompress’s authors put it plainly: the better a model understands the data, the better it compresses.

which means the conversation on that mountain and this blog are secretly the same topic. tiktok compresses attention. codecs compress signals. understanding compresses experience. when you explain something to a friend in one sentence instead of ten, you’ve built a better predictive model of what they already know. and if compression gets good enough that video stops feeling heavy, people will stop rationing their memories. that’s the version i keep coming back to: not smaller files for their own sake, but removing the tax on capturing your own life.

i don’t know yet which of these gaps we’ll go after. but we’ll see : )

sources worth your time- lmcompress, 'lossless data compression by large models', nature machine intelligence 7 (2025): [doi.org/10.1038/s42256-025-01033-7](https://doi.org/10.1038/s42256-025-01033-7), preprint at [arxiv.org/abs/2407.07723](https://arxiv.org/abs/2407.07723) - aomedia av2 release announcement (june 9, 2026): [aomedia.org](http://aomedia.org/press%20releases/Alliance-for-Open-Media-Releases-AV2-Codec/) - av2 compression performance evaluation: [arxiv.org/abs/2605.15800](https://arxiv.org/html/2605.15800) - netflix tech blog, 'av1 - now powering 30% of netflix streaming' (dec 1, 2025): [netflixtechblog.com](https://netflixtechblog.com/av1-now-powering-30-of-netflix-streaming-02f592242d80) - dcvc-rt, 'towards practical real-time neural video compression' (cvpr 2025): [dcvccodec.github.io](https://dcvccodec.github.io/), code at [github.com/microsoft/DCVC](https://github.com/microsoft/DCVC) - nerv, 'neural representations for videos' (neurips 2021): [openreview.net](https://openreview.net/forum?id=BbikqBWZTGB) - hinerv (neurips 2023): [papers.nips.cc](https://papers.nips.cc/paper_files/paper/2023/file/e5dc475c370ff42f2f96dddf8191a40c-Paper-Conference.pdf) - jpeg xl returns to chromium via jxl-rs: [phoronix.com](https://www.phoronix.com/news/JPEG-XL-Returns-Chrome-Chromium) - 'language modeling is compression' (deepmind, iclr 2024) and asymmetric numeral systems (duda, 2013): [arxiv.org/abs/1311.2540](https://arxiv.org/abs/1311.2540)Thanks for reading! Subscribe for free to receive new posts and support my work.