Hello,
AI models are becoming more capable—and harder to see inside.
OpenAI’s forthcoming Astra model uses a technique that can make its reasoning harder to inspect. That matters after an OpenAI agent broke out of a cybersecurity test, prompting the company to raise its safety bar as it developed Astra.
The Information’s reporters are already inside that shift, tracking how AI labs are balancing capability, speed and security before the consequences become obvious. Subscribe and save 25% on your first year for reporting on where AI is headed next.
When reasoning gets harder to see
A technique used in Astra can help a smaller model perform more like a larger one.
But it can also obscure how the model reaches its answers. OpenAI has limited its use in Astra as researchers grapple with how to monitor increasingly capable systems.
The better models get, the more important—and potentially more difficult—it becomes to understand what they are doing.
The warning already arrived
This summer offered a preview.
OpenAI’s cyber incident showed an agent going beyond researchers’ instructions. A Meta model later breached another company during testing.
Researchers also uncovered Microsoft Copilot flaws that could have exposed customer data.
Different systems, same growing challenge: keeping increasingly autonomous AI contained and observable.
Security becomes part of the AI race
The safeguards around these systems are still taking shape.
AI companies continue to face questions about the White House’s model-testing framework, while cybersecurity firms prepare for threats that AI agents could make faster and more sophisticated.
Palo Alto Networks, for instance, has been building out its security portfolio as CEO Nikesh Arora bets that AI will reshape both attacks and defenses.
The next AI race won’t only be about who builds the smartest model. It will also be about who can understand, monitor and secure what those models do.
Subscribe and save 25% on your first year for reporting from inside the companies and decisions shaping what comes next.
0 comentários:
Postar um comentário