Why we no longer evaluate SWE-bench Verified
OpenAI discontinues SWE-bench Verified due to contamination and training leakage; recommends SWE-bench Pro.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
OpenAI discontinues SWE-bench Verified due to contamination and training leakage; recommends SWE-bench Pro.
OpenAI launches Frontier Alliance Partners program to help enterprises deploy AI agents from pilots to production with secure, scalable infrastructure.
OpenAI submits proof attempts to First Proof math challenge, demonstrating research-grade reasoning on expert-level problems.
OpenAI commits $7.5M to The Alignment Project for independent AI alignment research addressing AGI safety and security.
OpenAI expands into India with local infrastructure, enterprise partnerships, and workforce development initiatives.
OpenAI presents Evals framework as tooling to bridge AI experimentation and production deployment, addressing the pilot-to-production gap enterprises face.
OpenAI and Paradigm introduce EVMbench, a benchmark for evaluating AI agents on smart contract vulnerability detection and exploitation.
GPT-5.2 proposes novel gluon amplitude formula in theoretical physics, later formally proved by OpenAI and academic collaborators.
OpenAI introduces Lockdown Mode and Elevated Risk labels in ChatGPT to defend against prompt injection and data exfiltration attacks.
OpenAI releases GABRIEL, an open-source toolkit using GPT to convert qualitative text and images into quantitative data for social science research.
OpenAI describes real-time access infrastructure combining rate limits, usage tracking, and credits for Sora and Codex.
OpenAI releases GPT-5.3-Codex-Spark, a real-time coding model with 15x faster generation and 128k context, in research preview.
OpenAI technical staff discuss engineering patterns for building agent systems with Codex as foundation model.
OpenAI tests advertising in ChatGPT free tier with privacy controls and answer independence guarantees.
OpenAI deploys custom ChatGPT instance on GenAI.mil for U.S. defense and intelligence operations.
OpenAI outlines localization strategy for adapting frontier models to regional languages, regulations, and cultural contexts.
OpenAI introduces trust-based access framework for deploying frontier cybersecurity capabilities with safeguards against misuse.
OpenAI Frontier is enterprise platform for building, deploying, and managing AI agents with governance and context management.
OpenAI releases Codex App Server with JSON-RPC API for streaming agent workflows, tool use, and code diffs.
OpenAI outlines Sora feed design philosophy emphasizing personalized recommendations and parental controls.
OpenAI and Snowflake announce $200M partnership embedding frontier models and agents directly in Snowflake's data platform.
OpenAI releases Codex app for macOS enabling parallel multi-agent coding workflows with long-running task support.
OpenAI disrupted Russia-linked influence operation using AI to generate comments on Russian cult leader arrest in Argentina.
OpenAI banned Rybar network accounts (likely Russia-origin) using AI for multilingual influence campaigns across websites and social platforms.
OpenAI banned Operation No Bell, a likely Russia-origin influence campaign using AI to produce anti-US criticism targeting African audiences.
OpenAI banned Cambodia-origin accounts using AI to conduct romance scams against Indonesian targets via translation and engagement automation.
OpenAI banned Cambodia-origin accounts using AI to impersonate recovery services and authorities, targeting fraud victims with false recovery schemes.
OpenAI banned accounts automating romance scam workflows using AI for outreach, translation, engagement, and investment fraud solicitation.
OpenAI banned likely China-origin accounts researching US persons and social-engineering tactics to enable targeted influence operations.
OpenAI banned China-linked accounts using AI for influence planning, harassment, and coordinated online operations.