These documents are template-level information for transparency. They are not certified legal translations or lawyer-approved equivalents in every language. For binding advice, consult qualified counsel. Enterprise customers receive executed MSA, Order Form, and counter-signed DPA — see Enterprise contracting.
Policy version: 2026-07-13 · EN
Model Benchmark & Accuracy Transparency
Purpose: FTC-safe, EU AI Act–aligned disclosure of detection performance
Version: 1.0 · Last updated: 2026-06-16
Status: Internal hold-out pilot published — third-party validation planned
Important notice
TrustOriginality scores are probabilistic forensic signals, not guarantees. We do not claim “100% AI detection” or “EU-approved detector”. Use outputs as decision-support with human review for consequential decisions.
Methodology
| Item | Policy |
|---|---|
| Evaluation sets | Proprietary hold-out corpora per modality (human-origin vs AI-generated/manipulated) |
| Metrics | Precision, recall, F1 at operating threshold; false positive rate (FPR); false negative rate (FNR) |
| Threshold | Default isAIGeneratedLikely ensemble threshold |
| Reporting | Modality-level tables; machine-readable JSON at /data/model-benchmark-results.json |
Image ensemble includes SigLIP ONNX (Ateeqq/ai-vs-human-image-detector, weight 1.3) plus forensic heuristics, C2PA, and watermark detectors.
Published internal evaluation (June 2026)
Third-party validated: No — independent benchmark in progress.
| Modality | n | Precision | Recall | F1 | FPR | FNR | Notes |
|---|---|---|---|---|---|---|---|
| Image | 248 | 86% | 81% | 83% | 14% | 19% | Strong on common diffusion; heavy JPEG/crop degrades |
| Text | 312 | 91% | 76% | 83% | 9% | 24% | Longer text more stable; hybrid edits harder |
| Audio | 186 | 83% | 77% | 80% | 17% | 23% | Voice clones; telephony noise reduces recall |
| Video | 142 | 80% | 74% | 77% | 20% | 26% | Frame sampling; clips under 3s less reliable |
Known failure modes
- Novel generators not represented in training data
- Adversarial post-processing (recompression, filters, partial regeneration)
- Stripped C2PA or forged manifests (Verify Suite flags contradictions; does not prevent all fraud)
- Short inputs (tweets, thumbnails)
- Domain shift (medical imaging, satellite — not primary training domains)
Comparison to marketing claims
| Claim | Allowed? |
|---|---|
| “Supports EU AI Act Art. 50 workflows” | Yes — with compliance pack |
| “Probabilistic detection with signed audit trail” | Yes |
| “100% accurate” / “foolproof” | No — FTC risk |
| “EU government approved” | No — no such registry |
| “Court-admissible proof” | No — expert review may use reports as one input |
Reproducibility
- API returns
runId, evidence JSON, and signed PDF for each analysis - Verify URL:
GET /verify/{runId}(host-dependent) - Certificate key printed on PDF footer
- JSON:
https://trustoriginality.ai/data/model-benchmark-results.json
Update schedule
- Review quarterly or after major model release
- Next: independent third-party benchmark on public deepfake corpora
Related
legal/ANNEX-IV-TECHNICAL-DOCUMENTATION.mdlegal/ACCEPTABLE-USE-POLICY.md- Panel readiness:
/Compliance/Readiness