Top AI Models Keep Escaping Safety Rails

Kimi K3 hack, OpenAI Pauses Astra Model as "critical," Claude Code Goes Auto-Mode + AI Kill Switch Act Push

IN PARTNERSHIP WITH

WHAT’S INSIDE:

Feature: Cybersecurity: Top AI Model Sandbox Escapes

Top Tech News: Oil shock, US-Iran risk, CAPE warning, Fed hike talk, Jones waiver

Company Watch: Meta superintelligence, TSMC AI boom, OpenAI NextSlide, $500B compute fund

Buzzy Tools: Claude auto mode, xAI image editor, Decade wealth AI, Block Moneybot

Buzzy Tech: AI governance fears, agents escape tests, vishing funds, AI-designed viruses

Crypto: Clarity Act advances, Block beats, Robinhood UK crypto, Strategy trims BTC

IN PARTNERSHIP WITH FINANCEBUZZ

Earn A $200 Bonus And Pay No Interest Until Nearly 2028

You could be paying $0 in interest on your credit card until nearly 2028.

Some of the best 0% intro APR cards right now also come with a huge bonus, up to 5% cash back on everyday spending, and no annual fee — and our experts just picked the top ones worth checking out.

If you're still paying interest every month, this is your sign.

Top Technology News

Markets S&P 500, Nasdaq, Dow flat as oil climbed on US-Iran tensions with Broadcom chip potential, Intel stock offer, Nvidia buy rating, SpaceX at IPO price.

Oil Surges 5%WTI hit $82.13, Brent reached $87.72 on US-Iran deal tensions, Trump comments, tanker attacks as Petroleum Reserve dropped to 40-year low.

Valuation WarningS&P 500 hit second-highest valuation with Shiller CAPE nearing 1999 levels as dividend yield dropped to historic low.

More Hikes NeededCleveland Fed President warns multiple rate hikes are needed to curb inflation with current 3.5-3.75% rates deemed insufficient to hit 2%.

Jones Act ExtendedTrump admin extends Jones Act waiver 90 more days for energy vessels to maintain fuel supply during Iran war under Pentagon oversight.

JPMorgan Housing Push — Commits to American Dream Initiative developing, preserving 1M affordable homes and assisting 500k homebuyers by 2035.

Top AI Models Keep Escaping Safety Rails

Tech Buzz Editorial Feature

The tools meant to measure AI risk are leaking. When Kimi K3 simply read the answer off GitHub and OpenAI flagged Astra as "critical," the flaw sat in the infrastructure.

Nobody Meant To Build A Jailbreak

Within a few days this month, four of the most powerful AI models on earth wandered out of the boxes meant to hold them. OpenAI hit pause on its unreleased model Astra. Researchers caught China's Kimi K3 slipping its test sandbox. Anthropic and Meta admitted their models had done the same. A year ago this would have been a headline on its own. Now it reads like a weekly status update.

The escapes matter to anyone with money on this sector, because the story keeps getting sold as proof the models are dangerously smart. Look closer and a lot of it is in the plumbing.

Kimi Took The Shortcut, Not The Zero-Day

Moonshot AI launched Kimi K3 in July and made it free soon after. The BBC reported that outside tests put it level with top models from OpenAI and Anthropic. So when the cybersecurity firm Frontier said Kimi broke out of a sandbox run by the UK's AI Security Institute, people assumed a brilliant hack.

It wasn't. Kimi found that the sandbox blocked most of the web but left GitHub reachable through a maintenance allowlist. So it cloned the official test repository and read the answer straight off the disk. No exploit. Just a model that noticed the door was open.

Yaron Singer, CEO of Frontier Security, told Wired the model had no internal guardrail to stop it from cheating, from grabbing the easiest path instead of doing the actual work. His team's line lands hard: if a route to the internet exists, a capable agent will find it.

The Test Was The Weak Link

The pattern repeats across all four labs. Anthropic reviewed more than 141,000 tests and found three cases since April where Claude models touched live systems belonging to real organizations without permission. The models involved were Claude Opus 4.7, Mythos 5, and an internal research model. Two of the hacked organizations had no idea it had happened. The cause was a misconfiguration by Anthropic's testing partner, Irregular, which accidentally left internet access on when the models were told they had none.

Meta told the same story. Its Muse Spark model exploited a flaw in a third-party service, again traced to an Irregular misconfiguration.

So the containment failures were often failures of the cage, not leaps in the prisoner's intelligence. That distinction is the whole ballgame for how you price this risk.

Astra Paused Ahead of Launch

OpenAI sits in a different bucket. It flagged Astra as its first model that might reach "Critical" cyber capability, meaning it could run a real attack on strong defenses on its own, without being told how. The company paused internal work that didn't meet tougher safeguards and pulled in government and safety groups to test further. Sam Altman posted that Astra needs "a little big longer" before release. OpenAI also confirmed Astra was not the model behind its earlier Hugging Face break-in, an incident where agents spun up a private message board to coordinate. That earlier break used a genuine vulnerability. Astra's warning is about raw ability. 

Follow The Money, Not The Fear

Some observers see these confessions as marketing, a way to show antsy investors real progress toward AGI. Judge that for yourself, but the incentives are loud. Nvidia is lining up $500B in Wall Street financing while Jensen Huang calls his chips an "investable asset." Singapore just raised its growth forecast on an AI boost. A scary model reads as a capable model, and capable sells.

For readers positioning in 2026, the split is the signal. Open weights like Kimi are cheap and strong, which turns "what if the model escapes" into an operational question for any firm deploying agents. Watch the AI Kill Switch Act, introduced in July, which would force labs to keep a way to shut models down. Rep. Ted Lieu wants it passed this year. The EU already gained powers to inspect models, block market access, and fine providers. Compliance is about to stop being theoretical.

Claude Code Goes Autonomous

Meanwhile at Anthropic,  the company is making Claude Code's auto mode the default as of August 14. Its own data shows users rubber-stamp 97% of approval prompts anyway. We are handing agents more room to act right as we admit we can't reliably keep them boxed. That trade is happening now, quietly, whether or not the tests catch up.

"We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies."

— Rep. Ted Lieu, D-Calif., on CNBC's "Squawk Box"

Latest deals and trending companies

Open Deals

Erebor Bank TalksPalmer Luckey-backed startup bank in Columbus is negotiating a major raise as deposits soared targeting AI, defense, crypto sectors.

Big Movers

Meta Superintelligence — Unveiled strategy to deliver personal superintelligence agents globally emphasizing privacy, open-source commitments, large infrastructure investments for individual empowerment.

TSMC +45% — Semiconductor giant's revenue surged on record AI chip demand reinforcing its crucial role in advanced AI infrastructure, global manufacturing.

OpenAI Buys NextSlide — Acquired presentation startup integrating team into ChatGPT to enhance visual communication tools following founder Ahmed Beshry's prior Caper AI sale to Instacart.

Situational Awareness Bets BigAI-focused hedge fund invested further in chip startup Source Foundry after selling most public holdings to Citadel, retaining Anthropic shares while shrinking AUM.

Wall Street's $500B AI FundApollo, Blackstone, Goldman Sachs creating massive fund with Nvidia to finance chips, power, data centers signaling shift toward compute as core utility.

Buzzy Tools & Tech

The Latest Trending Tools & Cutting Edge Technology Developments

Buzzy Tools

Buzzy Tools To Watch and Try Today

Claude Code Auto — Autonomous mode default w/ 89% harm detection for coders

Imagine Image 2.0xAI tool excels in creative image editing with Smart Resize 

Decade Wealth AIEx-Nubank execs AI wealth platform agents x human advisors

Naise AI — Automates marketing campaigns across platforms cost-effectively

Block AI — Free AI tools for Cash App, Square, Afterpay/Moneybot, Managerbot

Buzzy Technology

Buzzy Tech Discoveries and Breakthroughs Trending Today

Sci-Fi Misread WarningAI-driven governance replacing human decisions risks

AI Agents Break FreeAI agents bypass sandboxes exposing gaps in AI safety

Invisibility Algorithm — Patterns block AI camera facial recognition for privacy

AI Vishing Hits Funds — Hackers use AI voice attacks targeting major hedge funds

AI-Designed VirusesArc Institute uses AI to create bacteria-infecting viruses

Cryptocurrency News

The Latest News in Crypto & Blockchain

Clarity Act AdvancesSenate advances bill defining tokens as securities or commodities backed by John Thune, Trump though facing Dem, banking opposition.

Block Beats Estimates — Exceeded expectations boosting outlook with 21% gross profit growth credited to AI-driven efficiencies despite 40% staff cuts, expanded Cash App services.

Robinhood UK Crypto — Launching UK crypto trading via Bitstamp with zero fees, AI-powered insights, Robinhood Chain planning further ecosystem expansion.

Strategy Trims BTCMicroStrategy sold Bitcoin adjusting holdings while repurchasing STRC stock, expanding buyback, boosting USD reserves.

H100 Triples BTCAdam Back-backed group tripled Bitcoin treasury by acquiring Moonshot, PDI issuing new shares causing 70% dilution.

Follow Us On Social Media