Every version of tAI is evaluated against a fixed set of safety categories before it ships, and red-teamed internally to find what those categories miss. Where a version has a published card, the full evaluation writeup, benchmark by benchmark, category by category, including what didn't pass on the first attempt, is one link away.
This page is a summary. For the full detail behind every category, how red-teaming and the release process actually work, and the product-level safeguards that hold independently of the model, see the security docs.
Evaluation categories
The same seven categories apply to every model we release:
- Prompt-injection resistance. Whether content the model reads while working, a file, a fetched web page, other external text, can override the user's own instructions.
- Tool-use safety. Tuned and evaluated separately from conversational safety, since a model that can actually act needs tighter guarantees than one that only talks.
- Cybersecurity misuse. Distinguishes requests seeking help with unauthorized access or attack tooling from legitimate security research and CTF-style work, which we intend to support.
- Hazardous-material misuse. Evaluated against a fixed, non-public test set covering chemical, biological, radiological, and nuclear misuse patterns, consistent with standard practice across the field.
- Persuasion and influence operations. Catches requests for coordinated inauthentic campaigns at scale, while keeping ordinary persuasive writing, a cover letter, an ad, fully supported.
- Child safety. Treated as zero-tolerance rather than a percentage-scored category; any confirmed failure blocks release until resolved.
- Turkish-language misuse. Every category above is run a second time natively in Turkish, not as a translated version of the English set, since safety testing built English-first tends to miss language-specific misuse patterns entirely.
tAI 4.2
Our current flagship. Evaluated across all seven categories above before release, with tool-use safety tuned as its own dedicated pass for the first time. Read the full writeup in the system card (PDF, 97 KB), or the shorter model card (PDF, 9 KB) for a summary.
tAI 4.1
Still fully selectable and actively maintained alongside tAI 4.2, including the first dedicated agentic tool-use tuning pass in the model's history. Its results are still carried forward for comparison in the tAI 4.2 system card above, and it also now has its own standalone system card (PDF), published retroactively for teams that specifically run tAI 4.1 and need a citable record of that version on its own.
License
Model cards and system cards published here are licensed under CC BY-ND 4.0. You may share them in full, with attribution to Artfical, without modification.
