Public red-team reports now ship with capable models
Every model above our capability threshold now carries a public red-team report on its listing page — the full eval suite, plus links to the underlying eval traces.
By Safety & Trust Team

Physical-AI models act in the real world, so the bar for understanding what they can and can't do has to be higher than a README. Starting now, every model above our capability threshold ships with a public red-team report attached to its listing page.
What's in the report
- Results across the standard evaluation suite we run on capable models.
- Direct links to the underlying eval traces, so claims are auditable.
- Known failure modes and the conditions that trigger them.
- The acceptable-use boundaries that apply to the listing.
Why we're making it public
A safety report that only the platform can see doesn't build trust — it just moves the risk to buyers. Publishing the report, and the traces behind it, lets anyone evaluating a model check our work before they deploy it near people or hardware.
This sits alongside the rest of our safety program: the acceptable-use policy, abuse handling, and moderation review documented in the safety-and-moderation guide.


