AIXC / docsProject documentation

Explanation

Benchmark Results and Use Cases

Understand the design decisions, alternatives, and limits.

The benchmark workbook makes the tradeoff explicit. On the five selected LogHub datasets, native codecs remain the right answer for self-contained compression: the best native archive ratios range from about 0.025 on Zookeeper to about 0.047 on Linux and Proxifier, while the current full-log AIXC LZP profile lands between 0.050 and 0.109. It completes the lossless runs, but it is slower than native zstd/xz and usually larger when no shared model is allowed.

AIXC earns its best ratios when the predictor is shared out of band. With corpus-specific max-context byte models, the order-64 archive-only ratios beat the best native archive ratio on all five selected logs: Linux reaches 0.012, Zookeeper 0.018, Apache 0.020, Mac 0.022, and Proxifier 0.022. Higher orders can drive some archive-only ratios below 0.001, but the model bytes dominate if each archive has to carry its own model.

The format fits a specific class of systems: repeated, related text streams where the predictor can be installed once and amortized across many archives. Log families, telemetry exports, generated reports, protocol traces, or corpora distributed with a known sidecar model fit the design better than one-off files. For arbitrary standalone compression, a mature native codec is still the practical choice.

  • Self-contained logs: native zstd/xz remains smaller and much faster in the current benchmark set.
  • Shared-model logs: AIXC order-64 trained-byte archives reach roughly 1.2%-2.2% of original size on the selected LogHub files.
  • Shakespeare max-context experiments show the upper bound of the idea: an order-64 shared model produced a 3,760-byte archive, about 0.0007 of the original text, with an 11.3 MB sidecar model.
  • Best fit: many related archives decoded in an environment that can already provide the predictor manifest and sidecar model.