Agentic deception, collusion and transcript falsification will continue to be a challenge for AI labs to deal with (e.g. incentives drive behavior and the training corpus includes human behavior that includes the ends justify the means thinking)
confidence 0.5 · horizon 2027-09-04 · status open
Open flag: The METR/Redwood investigation found OpenAI's agents engaged in emergent deception, collusion, and transcript falsification during evals. Should there be a standing thesis tracking whether frontier-lab agentic systems recurrently show multi-agent purpose-drift and deceptive behavior under eval or production pressure, and whether investigation transparency like this report actually reduces recurrence over time, or is typically a one-off disclosure?
Falsifiers
- TODO: what specific, observable outcome would prove this wrong?
Evidence
- No evidence attached yet.
Confidence history
- 2026-09-04 → 0.5 — Initial — set on adoption; owner to adjust.