A Zero-Entropy System Has No Smooth Positive-Volume Model in Any Dimension
An OpenAI preprint constructs an ergodic, invertible, probability-preserving system with zero entropy that has no smooth model preserving a positive smooth volume density in any finite dimension. The authors’ argument turns on a mismatch: smooth dynamics would allow orbit pasts and futures to be spliced, while the system’s nested symbolic constraints force those halves to remain aligned. Zero information production per step, the paper shows, does not guarantee a smooth representation.

Zero entropy does not make a system smoothly representable
The paper constructs a single ergodic, invertible, probability-preserving system with zero entropy that has no smooth positive-volume model in any finite dimension. Zero entropy means that the long-run information produced per step vanishes; it does not mean that every observation is predictable or that the system has no structure. The obstruction lies in the constraints linking its past and future.†
A model here is an infinitely differentiable, invertible map on a compact manifold that preserves a probability measure with a smooth, strictly positive density. The symbolic system may be relabeled on invariant sets of full measure by a measure-preserving bijection, measurable in both directions. The relabeling need not be continuous, but it must respect the dynamics: relabeling after one symbolic step must agree with taking one smooth step after relabeling.
The theorem rules out such models in every finite dimension, including non-orientable manifolds and manifolds with smooth boundary. Its scope depends on the measure condition: it does not rule out smooth realizations with singular invariant measures, or establish the same obstruction for every non-smooth density.
Smooth dynamics would let the proof splice a past to a future
The smooth side of the contradiction begins with a switching operation. Partition the manifold into small cells, choose a cell with probability proportional to its volume, then draw two points, x and y, independently from the cell’s normalized volume. For most such pairs, the paper constructs a point z whose future stays close to x’s future and whose past stays close to y’s past, over a long finite interval.
In ideal product coordinates, this switch combines the slow coordinate of y with the fast coordinate of x. The diagram is only schematic: the actual construction corrects the ideal point to produce a genuine orbit. The important implication is that smoothness would allow the system to join two independently chosen orbit halves.
Closeness alone is not enough. The switched outputs must not concentrate in rare regions of the manifold. For every measurable set E, the original probability of retained pairs whose output lands in E is at most twice the volume of E. The retained pairs are not renormalized after discarding the others. This volume control lets the argument work even when the measure-preserving relabeling is only measurable.
The switching construction also requires economical partitions. At times n equal to 2 raised to the k, their logarithmic size is bounded by a constant times B-sub-k times n, where the sum of B-sub-k to the sixteenth power is finite. The paper derives these controls from derivative estimates and a Jacobian argument; the explainer omits those technical proofs.
The symbolic system hides constraints in the alignment of its words
To oppose smooth switching, the paper builds a symbolic system from nested legal words. Longer words are formed by concatenating shorter ones, with block boundaries recorded at each level. The words are separated enough that a sufficiently accurate copy reveals not just which word appears, but how it is aligned.
Each level assigns labels that are non-zero binary vectors. Certain pairs of labels, separated by specified powers of two numbers of blocks within the same group, must pass an orthogonality test. The construction must meet these local constraints while leaving short lists of labels with nearly maximal entropy. The paper proves that the requirements can coexist; the nesting diagram is organizational, not a computed example of the enormous word family.
A small version of the test uses four binary coordinates and a pairing computed modulo two: the first coordinate of one vector is multiplied by the third of the other, the second by the fourth, the third by the first, and the fourth by the second, with the products added. For the vector (1, 0, 0, 0), the pairing reads the third coordinate of the other vector. Eight of the sixteen possible vectors have that coordinate zero. Excluding the all-zero vector leaves seven passing partners among the fifteen non-zero possibilities; the fixed vector itself passes as well. This illustrates the relation, not the infinite construction. At larger dimensions, an entropy and matrix estimate controls independent label distributions.
Information per step vanishes, but the tests accumulate
Fix one word level, with word length m and L possible words. A window of n symbols can be specified by its starting phase, with m possibilities, and at most the ceiling of n divided by m, plus one, legal words. Thus the number of windows is at most m times L raised to the power of the ceiling of n divided by m, plus one. Taking the logarithm and dividing by n, then letting n grow at this fixed level, bounds the information rate by log base two of L divided by m. The construction makes these level rates tend to zero. The paper extends the count from phase observations to all finite measurable partitions.
That vanishing rate does not make the system’s tests harmless. At level i, write R-sub-i for the information rate and K-sub-i for the number of available test scales. Although R-sub-i tends to zero, the sum of K-sub-i times R-sub-i to the sixteenth power diverges. By contrast, smooth switching requires the sum of B-sub-k to the sixteenth power to converge.
The scale selection compares these two sums by grouping the symbolic tests into disjoint intervals of smooth scales. If every sufficiently late test in those intervals required B-sub-k to be at least a fixed multiple of its level’s R-sub-i, summing the sixteenth powers would make the smooth-cost sum diverge: the symbolic sum is already divergent. That contradicts the smooth estimate. So in arbitrarily late intervals, at least some selected tests have B-sub-k divided by R-sub-i arbitrarily small.
At those selected scales, learning which cell contains the pair removes only a negligible fraction of the label information. The cell label C costs at most the logarithm of the number of cells, which is negligible relative to the information in the label list. This is what lets the symbolic test retain its force even after conditioning on the cell.
One pair statistic yields incompatible bounds
The final contradiction compares two estimates of the same statistic under the same original pair distribution. Take a past list from y and a future list from x, then define Z as the fraction of corresponding label pairs that pass the orthogonality test. For the selected scales, conditional on the cell and a fixed entry, the two inputs are independent. Their remaining entropy implies that the expected score is at most 0.620.
If a smooth switch existed, its output would make the two lists closely match the past and future of one legal symbolic point. Word separation would recover a common alignment and most labels; volume control bounds the exceptions, including group boundaries. The paper’s lower bound on the expected score is 0.912. Since both estimates concern the same original law, they cannot both hold.
| Estimate for Z | Bound | Why |
|---|---|---|
| Independent inputs, conditional on the cell | At most 0.620 | Remaining label entropy |
| Lists matched by a legal switched point | At least 0.912 | Alignment recovery and volume control |
The order of choices makes this one system an obstruction to every finite-dimensional model. The symbolic system is fixed first; only then is a putative model considered, and the test scale is chosen late enough to absorb its constants. The full proofs of word existence and smooth switching remain in the manuscript. The finite examples illustrate pieces of the argument, not its general proof.