AIXC / docsProject documentation

Reference

Predictor Families and Practical Meaning

Look up syntax, contracts, layouts, algorithms, and exact behavior.

The repository includes several predictor families because the format question is broader than any one model. The toy backend is for tests, the LZP byte backend is a fast deterministic full-file baseline, trained-byte models explore shared corpus priors, neural-byte models test learned next-byte prediction, and Hugging Face token models test the higher-overhead language-model path.

These predictors have different economics. LZP adapts from the file itself and has no large sidecar, but it cannot know a corpus before seeing it. Trained-byte and neural predictors can be much stronger when the model is shared, but their model files must be counted unless the deployment environment already provides them. Hugging Face models add tokenizer and runtime reproducibility problems on top of model size.

  • Adaptive byte predictors are easiest to reproduce and benchmark.
  • Shared trained predictors can make archives tiny, but only after the model cost is amortized.
  • Neural predictors are useful research baselines, not currently the fastest or smallest path in the repo.
  • Token-model predictors require exact tokenizer round trips and pinned runtime behavior.