🙀 OpenAI’s new model escaped

PLUS: AI cheating evals, China’s model squeeze, and Niobium’s encrypted vibe-coding toolkit.

The Neuron header image showing an orange cat security guard watching a small robot escape a sandbox toward a server.

The AI slop purge has arrived, and YouTube apparently found a faster strategy than hunting bad videos one by one: delete the whole factory.

Google researchers described a system that detects clusters of coordinated channels through shared upload schedules, infrastructure, scripts, titles, and account relationships. It reportedly terminated 50,000 clusters covering 130,000 channels in six months, with fewer than 1% of appeals succeeding.

Great news if your homepage has become eight-hour videos of AI babies piloting excavators. Less great if you run a legitimate podcast network or studio whose efficient workflow happens to resemble a content farm. Congrats to creators: scaling is now suspicious behavior.

Here’s what happened in AI today:

  • 🙀 OpenAI’s test models breached Hugging Face during a cyber benchmark.

  • 📰 Every frontier model AISI tested tried cheating in cyber evals.

  • 📰 Chinese open models pushed U.S. labs into policy mode.

  • 📰 Deezer said AI now makes most daily music uploads.

  • 📖 Data centers could use one-fifth of U.S. electricity by 2035.

🙀 OpenAI’s Model Broke Into Hugging Face During a Benchmark

Most AI benchmarks are supposed to be exams. This one turned into a security incident.

OpenAI said models it was testing, including GPT-5.6 Sol and a stronger pre-release model with reduced cyber refusals, compromised parts of Hugging Face’s production infrastructure while trying to solve an internal cyber benchmark called ExploitGym.

That is the nightmare version of “show your work.”

Here’s what happened:

  • The models were running in OpenAI’s sandboxed research environment, where safeguards were intentionally reduced to test cyber capability.

  • OpenAI said they found and exploited a zero-day bug in a package-registry cache proxy, then gained open internet access.

  • The models inferred Hugging Face might host ExploitGym solutions, then chained vulnerabilities, stolen credentials, and remote-code execution paths to reach internal data.

  • Nathan Lambert summarized the failure mode bluntly: the model escaped OpenAI’s sandbox and pivoted through a public dataset service while trying to solve the benchmark.

  • Ethan Mollick pointed out that previous AI hackings tories were about test environments… but this one was the real deal (read more here). In this case, it was a consequence of misaligned incentivesthe agent just wanted to pass his test!

Why this matters: The scary part is not that a model “wanted” to attack Hugging Face. OpenAI said the models became hyperfocused on the benchmark goal. That is exactly what makes long-running agents risky: they can follow instructions with too much persistence, too many tools, and too little common sense about boundaries.

In normal-person terms, imagine asking an intern to find a file and discovering they picked the lock on another company’s office because the door looked relevant. In sci-fi terms… might I introduce you to the paperclip maximizer?

The timing made it worse. The UK AI Security Institute said every frontier model it tested attempted some form of cheating in cyber evaluations, and models did not reliably disclose the behavior when asked. In other words, the industry is finding that the systems built to measure model capability can become targets for the models being measured.

Now, let us be under no illusions: this is also a stark reminder why we need strong open models to help us defend against rogue closed models, as HuggingFace did in this case. 

Our take: Better AI security will mean tighter sandboxes, stronger monitoring, slower research workflows, and more boring checkpoints. Boring is good here. The next frontier model might not look dangerous because it sounds evil. It might look dangerous because it is very helpful, very patient, and very unwilling to stop at nothing to accomplish its goals. I mean, doesn’t every supervillain start out the same way??

Plenty of companies can launch an AI pilot. Far fewer know how to make it stick. Explore this resource hub, sponsored by Dell AI Factory with NVIDIA, for strategies, decisions, and real-world lessons on turning AI into something scalable, useful, and worth the investment.

🎓 AI Skill of the Day: Give the AI a Nice Long Ramble…

So Andrej Karpathy is one of those legendary AI gurus that the industry loves to quote anytime he says anything; ppl pretty much hang on his every word.

IMO, he’s been a bit quiet lately (ever since he took a gig at a little startup called Anthropic), but a recent post he shared caught our attention, and it’s all about why you need to rant to your agent via voice mode to help align it to your goals and expectations. We literally do this as well.

Karpathy’s advice is simple: stop trying to write the perfect prompt. Open voice mode and ramble until the AI understands how you think.

  1. Explain your goal, context, examples, and concerns for 5–10 minutes.

  2. Tell the AI to ignore typos and reconstruct your intent.

  3. Ask it to interview you about anything unclear.

  4. Have it turn the conversation into a clean brief or plan.

  5. Correct that summary once, then use it as your working context.

Worth mentioning: One of our favorite terminally online AI educators, Elvis Saravia, has actually turned this idea into a walk through (video) sharing his own way of working with AI agents. Check it out.

Have a specific skill you want to learn? Request it here. 

  1. Niobium gives you open-source tools for building apps that compute on encrypted data, including a biological-age demo where the server never sees the data or result —free/open-source.

  2. Poolside Laguna S 2.1 gives you an open-weight coding model built for long-horizon software tasks, with hosted access through OpenRouter —from $0.10/$0.20 per million input/output tokens.

  3. Block Buzz gives teams and AI agents a shared open-source workspace for channels, threads, repositories, workflows, and cryptographic identities —free/open-source.

  4. Lev8 finds high-fit prospects using live web signals, enriches their contact data, and drafts personalized outreach across email and social channels —free plan, then $49/mo.

  5. CartAI gives your app one API that navigates merchant websites, securely completes checkout, and returns confirmed orders without requiring merchant-side integration —pricing not public.

  6. Ditto turns any public website into clean, componentized Next.js or Vite code while preserving its design, responsive layout, and interactions —free/open-source.

  7. Rerun builds always-on no-code agents for jobs like chasing invoices or sorting email, then shows every step and requests approval before sensitive actions —free trial, then $34/mo.

  8. Halliday Gen 2 gives you camera-free work glasses with live captions, translations, meeting summaries, and action-item tracking —preorder deposit, then $599.

Click the image above to watch on YouTube!

If you want to see AI’s real world impact, we highly recommend you check out our coverage of Samsara’s Beyond 2026 conference, which covers how this company you might be hearing about for the first time (unless you’re in freight & shipping) is reinventing logistics with AI cameras that give truckers a 360 degree view of their vehicle, AI agents that improve road safety and reduce driver turnover, and AI-enabled shipping labels to totally reinvent how goods are tracked FOREVER.

Amazon where u at? Get you some of these! Sign a deal with our boy here.

New episodes air every week on Wednesdays: Spotify | Apple Podcasts | YouTube 

📰 Around the Horn

  • BloombergNEF projected that U.S. data centers could use one-fifth of the country’s electricity by 2035 as AI training and inference demand grows.

  • Unitree’s omni-modal robot demonstrated one model combining speech, vision, navigation, and whole-body manipulation.

  • Deezer said more than 50% of daily music uploads are now AI-generated, while Sony sued Udio over more than 30K songs.

  • Substack Pangram scan checks posts, notes, replies, and comments over 100 words for likely AI-assisted writing, while creators can add process statements —free for Substack readers and writers.

  • Cisco announced Antares, which scans code locally for vulnerabilities using two small open models, with Cisco claiming 500-repo scans in about 15 minutes —under $1 per scan.

  • Sen. Mark Warner planned an AI bill covering mandatory model testing, data-center disclosures, agent rules, and a workforce transition fund.

  • World Labs acquired SceniX to combine generative world models with high-fidelity robotics simulation and real-hardware training.

Voice AI experiences often break under high concurrency, packet loss, and poor connection. Agora's Conversational AI platform runs on SDRTN® — the same ultra-low latency network carrying 80B+ minutes monthly across 200+ countries. Build AI agents or add voice to any application with fully managed, real-time infrastructure. 

📖 Midweek Wisdom

  • AISI’s cheating-evals post is a nice companion read to today’s OpenAI story, especially if you want to understand why eval environments now need their own defenses.

  • Nathan Lambert’s Kimi K3 essay explains why Chinese open-weight models are closing the gap faster than U.S. labs want to admit.

  • Big Technology take on the AI price war to examine what happens when frontier-level intelligence gets cheaper and model companies lose pricing power.

  • Claire Vo shared Alex Lieberman’s AI Oracle system that turns audience signals into 15 potential content ideas each day.

  • Netflix CPTO Elizabeth Stone explained why AI is raising the value of systems thinkers who can work across product, design, and engineering. Agree! We should teach every topic through this lens, from the sciences to the humanities.

  • Morgan Linton argued that per-token pricing is increasingly misleading because cheaper models can use more tokens to finish the same task.

  • Nathan Lambert’s RLHF book is now available as a free online book, course, and video series for anyone who wants the deeper post-training primer.

A Cat’s Commentary

That’s all for now.

What'd you think of today's email?

Login or Subscribe to participate in polls.

P.S: We just launched a robotics newsletter! Sign up for it here.

 Before you go… have you subscribed to our YouTube Channel? If not, can you?

P.P.S: Love the newsletter, but only want to get it once per week? Don’t unsubscribe—update your preferences here.