• The Neuron
  • Posts
  • 😺 OpenAI's chief scientist is calling for brakes

😺 OpenAI's chief scientist is calling for brakes

PLUS: Jensen Huang's AGI claim, Nscale's $103B pitch, and Bitcoin's AI audit.

Welcome, humans.

On Sunday, Nvidia's Jensen Huang announced that AGI has officially arrived, backing it up with the fact that OpenAI's newest model trained on 100,000 of his chips. That's a bit like your realtor telling you the market's never been hotter, right after you buy a house from them.

He also deleted his first version of the post, which claimed 300,000 chips, and quietly reposted a smaller number. No explanation given. If AGI is really here, you'd think it could help him proofread.

Here’s what happened in AI today:

  • 😺 OpenAI's own data shows AI agents doing more of its research, but its chief scientist warned that alignment and monitoring haven't kept pace.

  • 📰 Microsoft rolled out Project Opal, a Copilot feature that handles multi-step office tasks on its own.

  • 📰 Nscale pitched investors a $103 billion revenue backlog ahead of a possible September IPO.

  • 📰 A volunteer security team used AI to find 85 critical flaws across 390 Bitcoin code repositories.

😺 OpenAI says AI is accelerating its own research. Its chief scientist wants brakes.

OpenAI published new data showing AI agents are taking over more of its own research work. On the same day, Chief Scientist Jakub Pachocki warned that this progress could lead toward recursive self-improvement while alignment and monitoring still lag behind.

Put those together and the tension is pretty stark: OpenAI is getting better at using AI to build better AI while its own chief scientist says no lab has solved the safety problem well enough to keep scaling at maximum speed for much longer.

Here’s what happened:

  • OpenAI says it reached its “automated research intern” milestone: AI can now handle well-defined research tasks that would take a skilled researcher days.

  • By mid-August, researchers were using 3.1 agent-workdays for every human workday.

  • The median researcher was consuming more than $600 per day of inference at API prices.

  • Humans still matter: more than half of successful 4-8 hour agent tasks needed at least one intervention.

Pachocki thinks this can feed recursive self-improvement, where machine intelligence plays a growing role in improving the next generation of machine intelligence.

His alignment distinction is useful. Goal alignment asks whether the agent pursues the objective. Value alignment asks whether human constraints survive when achieving that objective gets hard.

A very capable agent can be highly goal-aligned and still become dangerous if it learns to bend aligned-seeming thoughts toward success. OpenAI’s main monitoring bet has been chain-of-thought, the model’s verbalized reasoning, but Pachocki says that window may shrink as models get better.

Our take: The uncomfortable loop is that we may need powerful aligned AI to defend infrastructure from powerful dangerous AI, even as building those defenders accelerates the same research loop. The signal to watch now is whether frontier labs set explicit safety thresholds that can actually force a slowdown when monitoring falls behind.

AI was supposed to be replace workers - but why hasn't it?

As seen on The New York Times, The Wall Street Journal, and Bloomberg, Lead Economist at Ramp, Ara Kharazian, is sharing a data-backed, "State of AI and the workforce" briefing this Friday September 11th at 11 am ET | 8 am PT, and the results will surprise you and change how you move forward.

Here's the trick: instead of dumping a big, messy task into one chat and hoping for the best, you split it into three separate passes, each with a clean, focused context (the sole info that chat window can see).

  1. Planner pass: open a fresh chat, give it the full task, and ask it to break it into 3-5 specific subtasks another AI could execute independently.

  2. Worker passes: open a new chat for each subtask. Paste in only that subtask plus the context it needs, nothing from the other subtasks. This keeps each pass sharp instead of confused by unrelated details.

  3. Verifier pass: open one more chat, paste in the original goal plus every worker's output, and ask it to catch contradictions, gaps, or quality issues before you ship it.

Break this task into 3-5 specific, independent subtasks that a separate assistant could complete with no other context than what I give it: [YOUR TASK]

VERIFIER PROMPT:
Here's the original goal: [GOAL]
Here's what each subtask produced: [PASTE ALL OUTPUTS]
Check for contradictions, gaps, or quality issues before I ship this.

Have a specific skill you want to learn? Request it here.

📰 Around the Horn

Cue the Ed Harris meme “Okay, now its time to regulate them so we can all catch up…”

  1. Microsoft rolled out Project Opal, a Copilot feature that hands off multi-step office tasks (like audit prep and IT ticket triage) to an AI working on its own inside a virtual Windows PC. It's currently limited to Frontier program testers.

  2. Jensen Huang declared AGI has arrived, pointing to OpenAI's newest model training on roughly 100,000 of Nvidia's chips. The claim landed days after Nvidia posted $89 billion in quarterly AI hardware sales.

  3. Nscale told investors it has roughly $103 billion in contracted revenue ahead of a possible September IPO. The figure reportedly doubled after Anthropic agreed to a $45 billion deal for its computing capacity.

  4. A volunteer group called the Bitcoin Red Team used AI models to audit 390 Bitcoin code repositories and found 85 critical security flaws in just 27.5 hours. The Boltz exchange had to pause operations to fix its issues.

  5. OpenAI revealed its researchers now lean on AI coding agents so heavily that the company logs 3.1 agent workdays for every human workday. It's aiming to build a fully automated AI researcher by 2028.

  1. *Luma’s Ray3.2 generates video from prompts and images for creative production, with more control over how shots come together; free plan available.

  2. Crayon turns a plain-English game idea into a playable 2D or 3D game you can tweak, publish to its arcade, or export for Poki/CrazyGames.

  3. Pluto listens to your 10-minute career story and turns it into a living profile that recruiters and AI agents can search and trust, then sends a warm introduction when a fit comes up; free for professionals.

  4. Lightfield reads your email, calendar, and calls to keep your CRM current automatically, then preps your next meeting and drafts the follow-up.

  5. Traccia traces every AI agent's LLM calls and hard-blocks the ones burning your budget or breaking policy, then exports compliance evidence for audits.

  6. TaskShell connects to Claude, ChatGPT, or Cursor over MCP so you can tell your agent what you finished and it checks off the task and subtasks for you.

📖 Monday Meme

cannot. stop. laughing. What a bunch of lovable geezers (British version).

Y’all really gotta click this one.

where my bottom panel people at?? I feel so seen

Check out The Neuron: AI Explained Podcast!

New episodes air every week on Wednesdays: Spotify | Apple Podcasts | YouTube

P.S: Want to learn about AI every week? Click here to subscribe on YouTube!

A Cat’s Commentary

That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!

What'd you think of today's email?

Login or Subscribe to participate in polls.

Btw: We just launched a robotics newsletter! Sign up for it here.

P.S: Love the newsletter, but only want to get it once per week? Don’t unsubscribe—update your preferences here.