2026年9月10日

Savage Move: Altman Personally “Shuts Down” GPT-5.2 as OpenAI Unveils Its Most Powerful Coding AI Yet

GPT-5.2-Codex: A Midnight Drop OpenAI has just launched GPT-5.2-Codex—its most powerful agentic codi...

GPT-5.2-Codex: A Midnight Drop

OpenAI has just launched GPT-5.2-Codex—its most powerful agentic coding model to date, built specifically for complex, real-world software engineering.

As the name suggests, GPT-5.2-Codex is an optimized evolution of GPT-5.2, with major upgrades across several key areas:

It introduces stronger context compression, making it better at handling long-running tasks without losing the thread.

It performs more reliably on large-scale code changes, including refactors and migrations.

Its coding performance in native Windows environments is significantly improved.

And OpenAI emphasizes that it reaches their highest level of cybersecurity capability so far.

Sam Altman says OpenAI teams are already using it internally—and seeing strong results.

In benchmark evaluations, GPT-5.2-Codex outperforms 5.1-Codex-Max, GPT-5.2, and GPT-5.1 on software engineering and terminal-based testing tasks.

OpenAI also highlights security performance repeatedly in its communications. In one recent example, a security researcher using GPT-5.1-Codex-Max with Codex CLI reportedly identified a React issue that could lead to source code exposure—showcasing how these tools can accelerate defensive security work.

Starting today, all paid users can access GPT-5.2-Codex, and the API is expected to open in the coming weeks.

Long-Run Coding Power That Doesn’t Drop the Thread

At its core, GPT-5.2-Codex is positioned as a “best of both worlds” release.

It retains GPT-5.2’s strength in professional task handling, while incorporating 5.1-Codex-Max’s improvements in agentic programming and terminal workflows.

The result is meaningful progress in long-context understanding, tool use, factual reliability, and native context compression—helping the model stay stable over long sessions while using tokens more efficiently.

In industry benchmarks, GPT-5.2-Codex is described as achieving state-of-the-art results on SWE-Bench Pro and Terminal-Bench 2.0, with an approximate 6% improvement over 5.1-Codex.

These benchmarks aim to measure how effectively models operate as agents in realistic terminal environments across diverse tasks.

Its Windows-native agentic performance also improves further, extending capabilities introduced in 5.1-Codex-Max.

With these upgrades, Codex is designed to work longer inside large codebases while maintaining full context—making it more dependable for demanding tasks like major refactors, migrations, and feature development, even when plans change midstream or early attempts don’t work out.

Better “Vision” for Real Developer Inputs

GPT-5.2-Codex also improves multimodal understanding.

Developers can share screenshots, diagrams, charts, and UI surfaces—and the model is intended to interpret them more accurately.

It can even read design mockups and rapidly translate them into runnable prototypes, which teams can iteratively refine until they’re ready for production.

Three Leaps Toward Real-World Capability

OpenAI describes a clear pattern of step-changes in a core cybersecurity evaluation:

GPT-5-Codex delivered the first major jump.

GPT-5.1-Codex-Max delivered the second.

GPT-5.2-Codex marks the third leap.

OpenAI expects this trajectory to continue, and notes that their planning assumes future generations may approach the “high” cybersecurity capability tier defined in their Preparedness Framework—though GPT-5.2-Codex has not reached that level yet.

A Real-World Security Case: From Disclosure to Discovery

On December 11, the React team disclosed three security issues related to React Server Components.

Andrew MacPherson, Chief Security Engineer at Privy (a Stripe company), used the situation to test how capable modern AI coding agents really are. While reproducing and studying the reported issues with GPT-5.1-Codex-Max and Codex CLI (alongside other agents), he reportedly uncovered an additional critical React weakness during the investigation, which was then responsibly reported to the React team.

The takeaway: advanced AI systems can meaningfully speed up defensive security research on widely used software—when used carefully and responsibly.

Early Community Tests Are Mixed

Some developers report failures on certain simulation-style coding tasks.

Others highlight strong results in areas like polished animations comparable to leading competitors.

There are also reports of impressive performance building game-like prototypes.

Overall, OpenAI frames GPT-5.2-Codex as a major step forward for real-world development and cybersecurity—helping developers tackle complex, time-consuming work more easily, while also strengthening tools available for defensive security research.

接著讀

In 49 Days, Brentford Emptied Out: Head Coach, Captain & Two Core Stars All Depart, £100M in Sales Won’t Ease Relegation Woes

On July 22 (Beijing time), Brentford confirmed the transfer of star forward Yoane Wissa to Manchester United for £71 million. In just 49 days, the Premier League’s surprise package—who finished 10th last season—has seen its head coach, captain, goalkeeper, and two key attackers all leave, raking in nearly £108 million but facing a brutal relegation battle next term.

415 天前