How we fact-check specifications, verify benchmark assertions, and maintain our index free from marketing hype.
We extract parameter counts, structural features, and context lengths exclusively from official release documentation, whitepapers, codebases, or directly from API response headers. We ignore third-party leaks and rumors.
Open-weights models are evaluated against OSI definitions. If a license imposes custom commercial restrictions (e.g., LLaMA, Mistral Research, or DeepSeek limits), it is categorized with strict clarity so commercial creators stay legally secure.
Coding and reasoning evaluations (e.g., SWE-Bench, Aider Polyglot, and GPQA) are cataloged only when published alongside structured code repositories or verifiable logs. Self-reported marketing evaluations are excluded from row-highlights.
Our automated monitoring engines parse academic papers on arXiv, index cards on Hugging Face, and developer announcement channels daily to register new model assets the moment they land.
A curator reviews the newly listed model's architectural specifications, verifying model type (open weights vs. API), context capacities, parameter sizes, and structural layers to ensure database integrity.
All extracted specifications are structured into our TypeScript-verified schemas. The compilation suite validates pricing specifications, task modalities, and parameter tags to maintain strict consistency.
Validated model models compile statically into pre-rendered routes, ensuring the comparison lists and indices remain lightning-fast and highly secure for developers referencing them daily.