Policy version: 2026-07-13 · EN

Model Benchmark & Accuracy Transparency

Purpose: FTC-safe, EU AI Act–aligned disclosure of detection performance
Version: 1.0 · Last updated: 2026-06-16
Status: Internal hold-out pilot published — third-party validation planned


Important notice

TrustOriginality scores are probabilistic forensic signals, not guarantees. We do not claim “100% AI detection” or “EU-approved detector”. Use outputs as decision-support with human review for consequential decisions.

Methodology

Item Policy
Evaluation sets Proprietary hold-out corpora per modality (human-origin vs AI-generated/manipulated)
Metrics Precision, recall, F1 at operating threshold; false positive rate (FPR); false negative rate (FNR)
Threshold Default isAIGeneratedLikely ensemble threshold
Reporting Modality-level tables; machine-readable JSON at /data/model-benchmark-results.json

Image ensemble includes SigLIP ONNX (Ateeqq/ai-vs-human-image-detector, weight 1.3) plus forensic heuristics, C2PA, and watermark detectors.

Published internal evaluation (June 2026)

Third-party validated: No — independent benchmark in progress.

Modality n Precision Recall F1 FPR FNR Notes
Image 248 86% 81% 83% 14% 19% Strong on common diffusion; heavy JPEG/crop degrades
Text 312 91% 76% 83% 9% 24% Longer text more stable; hybrid edits harder
Audio 186 83% 77% 80% 17% 23% Voice clones; telephony noise reduces recall
Video 142 80% 74% 77% 20% 26% Frame sampling; clips under 3s less reliable

Known failure modes

  1. Novel generators not represented in training data
  2. Adversarial post-processing (recompression, filters, partial regeneration)
  3. Stripped C2PA or forged manifests (Verify Suite flags contradictions; does not prevent all fraud)
  4. Short inputs (tweets, thumbnails)
  5. Domain shift (medical imaging, satellite — not primary training domains)

Comparison to marketing claims

Claim Allowed?
“Supports EU AI Act Art. 50 workflows” Yes — with compliance pack
“Probabilistic detection with signed audit trail” Yes
“100% accurate” / “foolproof” No — FTC risk
“EU government approved” No — no such registry
“Court-admissible proof” No — expert review may use reports as one input

Reproducibility

  • API returns runId, evidence JSON, and signed PDF for each analysis
  • Verify URL: GET /verify/{runId} (host-dependent)
  • Certificate key printed on PDF footer
  • JSON: https://trustoriginality.ai/data/model-benchmark-results.json

Update schedule

  • Review quarterly or after major model release
  • Next: independent third-party benchmark on public deepfake corpora
  • legal/ANNEX-IV-TECHNICAL-DOCUMENTATION.md
  • legal/ACCEPTABLE-USE-POLICY.md
  • Panel readiness: /Compliance/Readiness