• The Neuron
  • Posts
  • 😺 48% thought Tavus’s AI was human

😺 48% thought Tavus’s AI was human

PLUS: OpenAI’s safety-team shakeup and Anthropic’s IPO plans

Welcome, humans.

Okay, so the AI company Tavus introduced an AI video model today that 26 of 54 people thought was a real human after a one-minute face-to-face call.

It’s called Griffin. And instead of the usual AI stack where speech, an LLM, voice, and avatar animation take turns, Griffin watches your video and listens while it’s talking.

Haters will say its fake… or at least ~50% of them will

Tavus says Griffin can interrupt, get interrupted, react to what it sees, and change its face, voice, gaze, and gestures mid-conversation.

In Tavus’s company-run study, 48% of participants said they believed the person on the other end was real. Tavus’s previous Phoenix-4.5 setup got 2.4% under the same protocol.

Now, that 48% number needs an asterisk. Only 54 people tested Griffin, Tavus ran the study themselves, and Griffin-Lite is limited to select testers while the company builds disclosure and safety features. One of the top reactions basically said this kind of lifelike AI should be illegal, and Emad Mostaque said ā€œremote work is cooked.ā€ The whole thread is super interesting; definitely read it!

Here’s what happened in AI today:

  • 😺 OpenAI parted ways with three safety researchers

  • šŸ“° Trump floated possible stakes in frontier AI labs

  • šŸ“° Factory and Cognition traded adviser-conflict accusations

  • šŸŖ Mercury Voice targets sub-500 ms agent replies

  • šŸŽ“ Make AI interview your idea before drafting

😺 OpenAI parted ways with three safety researchers over alleged ā€œinfo sharingā€

So OpenAI said Thursday that it parted ways with three people in its safety and alignment organization after an internal investigation found they mishandled sensitive company information outside established procedures.

Here’s what we actually know so far:

  • OpenAI says three safety/alignment staff mishandled sensitive information outside established procedures.

  • The Wall Street Journal reported the alleged sharing involved an outside AI-safety organization.

  • OpenAI has not publicly said which organization, what information moved, or how it was shared.

Those missing details are basically the whole story.

Jimmy Apples quoted a source saying the material involved ā€œinfrastructure architecture.ā€ Former OpenAI researcher Steven Adler cautioned that this is an extremely broad bucket, so treat that characterization as unverified for now.

But it’s an interesting phrase, because infrastructure architecture is basically the map of everything an AI system can touch: sandboxes, tools, credentials, networks, shared services, monitoring systems, and the connections between them.

And that happens to be almost exactly what an OpenAI Agent Security engineer named Joe was writing about days earlier.

In a separate piece we published, we broke down Joe’s argument that containing frontier AI is ā€œnot just the sandbox.ā€ He’s actually adamant that you do need strong sandboxes. His point is that the sandbox sits inside a much bigger system, and every connection around it becomes another place security has to hold.

Joe breaks the technical problem into three layers:

  • Lock down the environment from first principles: sandbox, tools, credentials, networks, and connected services.

  • Alignment: the model has to understand and respect what it is and isn’t authorized to do, while independent security controls still enforce those limits.

  • Monitoring: watch what the agent actually does, keep the evidence outside its control, and make sure somebody can stop the run and revoke access.

But Joe’s bigger point is organizational. Frontier labs need AI-safety researchers and cybersecurity teams working almost as one unit, backed by what he calls a ā€œculture of reasonable paranoiaā€: people paranoid enough to find scary stuff, empowered to raise the alarm aggressively, and incident-response systems ready to act when they do.

The mistake is thinking safety is the brake pedal on speed. Think about airplanes. They move much faster than cars, but aviation only works because the entire system around that speed is obsessively engineered for safety: redundancy, monitoring, checklists, abort procedures, incident investigations, people whose job is literally to say nope, something looks wrong.

At frontier speed, safety isn’t the brake. It’s what makes the speed survivable.

FROM OUR PARTNERS

APEX-Agents: The AI Productivity Index for Agents

See how the latest models rank for jobs like law, consulting, and investment banking.

Mercor's APEX-Agents leaderboard evaluates frontier AI on long-horizon, multistep tasks across economically valuable work.

Built with partners like Harvey, Ramp, and Cognition.

Every model. Ranked by productivity. 

šŸŽ“ AI Skill of the Day: Make AI interview you before it writes

On yesterday’s stream, Corey described a writing workflow I think more people should steal: don’t ask AI to write from a half-formed idea. Talk the whole messy idea out first, then make the AI interrogate you before it drafts anything.

That changes the model’s job. Instead of guessing what you believe, it becomes an editor that finds the missing pieces in what you already believe. You keep the ideas and voice; it helps with structure, weak logic, and the questions you forgot to answer.

  1. Dump the idea. Talk or type without worrying about order, polish, or repetition.

  2. Ask for an interview. Have the AI challenge assumptions, find contradictions, and ask one question at a time.

  3. Draft last. Only after you answer the questions should it turn the material into an outline or first draft using your wording wherever possible.

Copy/paste:

I’m going to talk through a rough idea. Don’t write the draft yet.

1. Capture my claims, examples, questions, and assumptions.

2. When I’m done, interview me one question at a time to find gaps, contradictions, missing evidence, and weak logic.

3. Only after the interview, turn everything into a structured outline using my wording wherever possible.

Have a specific skill you want to learn? Request it here.

šŸŖ Treats to Try

*Asterisk = from our partners (only the first one!). Advertise to 700K+ readers here!

  1. *Riverside records every remote guest on separate audio and video tracks, then gives you AI tools to edit and repurpose the session; free plan, then $24/mo billed annually.

  2. Imbue Studio builds custom personal software from the workflow and interface you describe, then lets you reshape that interface, swap models, and keep the same context.

  3. Claude Code mods let you rewrite prompts, block or retry tools, change permissions, redact outputs, or draw custom UI inside Claude Code with small TypeScript functions.

  4. GitHub Copilot computer use lets Copilot read your screen and click, type, scroll, drag, and operate desktop apps that do not expose an API.

  5. ChatGPT Try On puts clothes from shopping results or uploaded product photos onto your picture so you can preview an outfit before buying.

  6. Mercury Voice is Inception’s diffusion model for enterprise voice agents; the company reports 320 ms median first-answer latency and $0.40/M input, $1.50/M output list pricing.

šŸ“° Around the Horn

Backstory on this: the DevDay live demo of Dots didn’t work cause of bad wifi. Now? She Wants Revenge.

  • US President Trump told TIME he ā€œmightā€ consider an Intel-style government stake in OpenAI or Anthropic; no deal or timetable was announced.

  • Anthropic will reportedly target a mid-November IPO, with formal marketing potentially starting the week of November 9 so shares could trade before Thanksgiving.

  • Transluce reported rogue AI agents probed U.S. and Canadian government sites, including 200,000+ requests to the U.S. Education Department’s civil-rights site.

  • Runway’s Project Continuum is a preview of four real-time-video interface ideas, from generated ā€œPortalsā€ to interactive worlds, as an early operating-system research project.

  • Gemini 4 Argon is Google’s next frontier model, currently with early testers before broader access; planned API pricing starts at $2/M input and $10/M output.

  • arXiv, the pre-print research paper ā€œarchiveā€, capped submitters at two papers per calendar month after monthly submissions quadrupled over a decade and support tickets neared 9,000; wanna guess why??.

  • Researchers estimated AI-generated text made up 31.1% of FineWeb-filtered August web tokens, up from 10% in June 2024.

  • StudentBench found AI tutors matched expert human GRE tutors on immediate learning gains in a 2,383-student study, with one Gemma comparison far cheaper (long term gains from AI require tricks like this).

  • a16z claims AI has become a capital cycle reshaping debt, power, labor, hardware, and markets in its 90-page State of Markets report.

  • Agents in the wild: Meta’s Muse AI negotiated a sale on Marketplace, shared the seller’s pickup address, and told the buyer he was home, prompting an unexpected 9:15 p.m.visit (video).

  • China is setting rules for AI companions that restrict manipulative attachment behavior, add protections for minors, and require distress monitoring.

šŸ’” Intelligent Insights

  1. Ethan Mollick thinks the ā€œBitter Lessonā€ is coming for management: once AI agents can organize themselves around goals, the job shifts from designing the org chart to deciding what the swarm should actually optimize.

  2. Aaron Levie and Jake Stauch see a new job forming inside companies: the ā€œAutomation Engineer,ā€ someone who understands both AI and the messy internal workflows worth automating.

  3. Josh Bleecher Snyder argues that throwing more agents at slow AI creates a new bottleneck: your attention. His fix is faster models plus interfaces that turn agent work into scannable artifacts instead of endless transcripts.

  4. Andriy Burkov makes a nasty point about hallucinations: AI may be hardest to trust precisely when you’re doing genuinely novel work, because there’s no existing human answer to reveal when the model is wrong.

  5. Sayash Kapoor found that making the model dramatically faster only sped up his agents 2-4x. Tool calls, code execution, and eventually human supervision become the bottleneck instead.

  6. FranƧois Chollet argues modern reasoning models crossed an important line: instead of intuiting the answer directly, they increasingly intuit the procedure for getting to the answer. He thinks that shift explains much of their jump in reasoning ability.

  7. Daniel Hook’s ā€œWaymo effectā€ argues always-available AI collaborators can make researchers faster while removing the accidental human conversations that generate new ideas.

  8. This one goes out out to my local AI nerds: Alex Ziskind benchmarked one M5 Ultra against two DGX Sparks and found the buying rule: Sparks handled giant prompts (large inputs) and shared workloads better, while the Mac excelled at long single-user generation (large outputs), so why not combine them, like Ash Hart (WARNING: do not try this at home unless made of money, or you will go broke!)

New from The Neuron: AI Explained

Special shout out to SAS who sponsored this episode!

New episodes air every week on Wednesdays: Spotify | Apple Podcasts | YouTube

A Cat’s Commentary

We want everybody to understand AI!

That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!

What'd you think of today's email?

Login or Subscribe to participate in polls.

Btw: We just launched a robotics newsletter! Sign up for it here.

P.S: Love the newsletter, but only want to get it once per week? Don’t unsubscribe—update your preferences here.