
Trade-offs
Bitrate against resolution, and something has to give
Where the shortage shows. A fixed bitrate spread over more pixels leaves less for each one of them.
Every setting is a trade · Long
Every encode is a budget. Spread it over more pixels and each one gets less — the question is only where the deficit shows.
The arithmetic is simple; the consequences are not
A video encoder is, at its core, an accountant working under a hard limit. The bitrate — the number of bits per second the finished file may consume — is fixed before the first frame is analysed. Everything the encoder does from that point is allocation: which parts of the picture earn more bits and which get by with fewer. Resolution determines how many mouths need feeding. Double the width and height, and the pixel count quadruples; the same bitrate budget now stretches four times as far. Stretch anything four times and something tears.
What tears, and where, depends on the codec's decisions. Early MPEG-4 encoders — including the hacked Microsoft build that became DivX ;-) 3.11 Alpha — had relatively crude allocation tools. The encoder divided the frame into macroblocks, typically sixteen by sixteen pixels, and decided for each block how much quantisation to apply. Quantisation is the lossy step: it rounds the frequency coefficients that represent the block's detail to coarser values, discarding fine information in exchange for fewer bits. Turn it up and the image softens; turn it up further and the block boundaries become visible as the blocky artifacting that defined the era. At the bitrates those early encodes were targeting — squeezing a feature film onto a 700-megabyte CD-R — quantisation was severe, and it showed on anything with fine texture: hair, grass, fabric.
What the encoder is actually choosing between
The fundamental tradeoff is not simply resolution versus quality; it is resolution versus detail density. A 640×480 encode at a given bitrate might look clean because each macroblock has enough bits to capture the local frequencies accurately. Upscale that same file to fit a larger screen and the encoder's approximations become visible at normal viewing distance. Encode the same source at 1280×720 at the same bitrate, and the encoder must either skip more spatial detail or skip more temporal detail — smearing motion between frames rather than describing it precisely.
Temporal decisions are where the keyframe interval and group-of-pictures (GOP) structure become significant. A keyframe stores a full frame independently; the frames between it and the next keyframe store only differences from their predecessors. At low bitrates, a scene with fast motion demands more bits for those difference frames because the differences are large. An encoder that cannot afford them either increases quantisation, degrading the visible frames, or drops frames, degrading motion smoothness. The first failure mode looks soft and blocky; the second looks stuttery. Both are a response to the same shortage.

XviD, the open-source MPEG-4 Part 2 encoder developed after DivX had taken the commercial path, exposed these choices more directly than its predecessors. Its two-pass encoding mode — where the first pass analysed the source for complexity and the second used that analysis to allocate bits dynamically — was an attempt to solve the uniformity problem. Simple static scenes got fewer bits, which were held in reserve for cuts, camera pans and sequences with fine detail. The bitrate in any given second could swing considerably above or below the target, provided the average across the file landed on budget. This was a meaningful improvement over constant bitrate encoding, where the encoder had no foresight and either wasted bits on easy passages or ran short on hard ones.
The container had to support this, and not all of them did cleanly. AVI was designed around the assumption of a roughly constant data rate — a legacy of its 1992 origins in the CD-ROM era, when drives required a predictable stream. Variable bitrate MPEG-4 inside AVI was technically workable but relied on workarounds; the index structures did not anticipate the kind of variation a good two-pass encode wanted to generate. Later containers, particularly Matroska, were designed from the ground up to handle variable-bitrate streams without such friction.
The generation problem, and what H.264 changed
Every time a lossy encode is re-encoded, the deficit compounds. The artefacts baked into the first encode — the softened edges, the smeared chroma, the quantisation noise — are treated by the second encoder as signal, not as noise, which means it tries to represent them faithfully and spends bits doing so, with worse results than the first pass. This generation loss made resolution decisions feel consequential in a way that lossless intermediate workflows do not. Encoding at too high a resolution for the bitrate budget, then later re-encoding for redistribution, produced files where multiple sets of artefacts overlapped: the original compression noise plus the noise introduced trying to approximate it.
H.264/AVC, standardised by the Moving Picture Experts Group and the ITU-T in 2003, redrew the tradeoff significantly. Its variable block-size motion compensation meant the encoder could use different block sizes in different parts of the frame — large blocks for simple areas, small blocks for complex ones — allocating bits at finer granularity than the fixed-macroblock MPEG-4 Part 2 approach. At equivalent bitrates, H.264 consistently delivered better results, which effectively meant that for the same perceptible quality, it needed fewer bits per pixel. That is the compression efficiency improvement that made 1080p practical at streaming bitrates.
How it works
The allocation chain
- Bitratethe per-second bit budget; fixed before encoding begins
- Quantisationthe step that rounds frequency data coarser; the primary lever on spatial quality
- Two-pass encodingfirst pass analyses complexity; second pass distributes bits accordingly; XviD popularised it in the open-source MPEG-4 era
- GOP / group of picturesthe span between keyframes; longer GOPs save bits on static material, cost more during motion
- Chroma subsampling (4:2:0)colour stored at half luminance resolution; standard across formats from DVD to streaming; exploits eye's lower colour acuity
The efficiency race continued. HEVC pushed further but entangled itself in patent pools. The Alliance for Open Media's AV1 was built explicitly as a royalty-free answer to that impasse, trading encode complexity — AV1 encoders are significantly more computationally demanding — for both improved compression efficiency and freedom from licensing cost. In each generation, the same underlying question repeated: given a fixed bitrate, what does the codec choose to preserve, and what does it let go?
Chroma subsampling sits underneath all of it, mostly invisible. Human vision is more sensitive to luminance variation than to colour variation, so encoders routinely store colour information at half the horizontal and vertical resolution of the brightness channel — the 4:2:0 scheme that became standard across broadcast, disc and streaming. It is a perceptual shortcut that works so well most viewers never notice it, which is the point. The bits saved on chroma are bits available for the luminance detail that the eye will actually scrutinise.

The arithmetic has not changed since the first MPEG-4 hacks circulated in the late 1990s. Bitrate is finite, pixels are hungry, and the encoder's job is to make the shortage invisible. Four generations of codec have refined the answer without altering the question.
How it works
Codec generations at a glance
- MPEG-4 Part 2 (DivX / XviD era, late 1990s–mid 2000s)fixed 16×16 macroblock structure; quantisation artefacts visible at low bitrates
- H.264/AVC (standardised 2003)variable block sizes; roughly doubled compression efficiency over MPEG-4 Part 2 at comparable quality
- HEVC/H.265 (standardised 2013)further efficiency gains; adoption slowed by fragmented patent pool licensing
- AV1 (published 2018 by the Alliance for Open Media)royalty-free; competitive with HEVC; substantially higher encode-time cost