research

Evaluation Reports

Published results from running the InclusiveCode eval suite against production LLM models. Each report documents pass rates, failure analysis, and safety gaps.

Want to run the eval suite against a different model?

Run your own eval →