- The Neuron
- Posts
- 😸 Multimodal AI just got real
😸 Multimodal AI just got real
PLUS: Anduril's $100B raise

Welcome, humans.
A politician appears to have read an AI-written speech aloud without removing the assistant’s offer to compile it as a PDF. The remarks apparently went from public policy to customer support in one uninterrupted breath.
AI can research the topic, draft the speech, and polish every sentence. It still cannot stop you from walking to the microphone without reading the document first. Human oversight has officially been reduced to “scroll to the bottom.”
Here’s what happened in AI today:
😻 Black Forest Labs unveiled FLUX 3 for video, audio, and robotics.
📰 NVIDIA, Microsoft, and Meta opposed broad open-weight restrictions.
📰 OpenAI reportedly delayed disclosing its role in the Hugging Face hack.
🍪 Handoff H1 automated construction takeoffs from uploaded blueprints.
🎓 Use 25-minute agent reviews to protect focus and judgment.
Hey: Want to reach 700,000+ AI-hungry readers? Advertise with us!
P.S: We just launched a robotics newsletter! Sign up for it here.

😻 FLUX 3 Can Make Videos, Generate Sound, and Control Robots
Most AI image generators are learning to make prettier pictures. Black Forest Labs is teaching its next model how the world moves, sounds, and responds when something touches it.
FLUX 3 is the company’s new multimodal foundation model, trained across images, video, audio, and robot actions inside one system. The pitch is bigger than creative software: learn enough about reality to generate it convincingly, then use that same understanding to act inside it.
Here’s what happened:
FLUX 3 can generate videos with native audio up to 20 seconds from text, images, existing video, or keyframes.
It supports multilingual dialogue, synchronized sound effects, animated text, multiple aspect ratios, and chained clips for longer sequences.
The same model backbone powers FLUX-mimic, a robotics system tested on Audi production tasks involving cables, seals, parts, and other objects traditional robots struggle to handle.
Early internal evaluations favored FLUX 3 over several rival video models, including Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%. Those results are preliminary, company-run, and likely to change as the system develops.
How to try it:
Open the FLUX 3 announcement.
Choose the early-access link for FLUX 3 Video.
Request access while BFL rolls availability out in batches.
Why this matters: Today’s creative models usually treat image, video, sound, and robotics as separate jobs. FLUX 3 treats them as different views of the same event. A ball hitting the floor has an appearance, a motion, a sound, and a physical consequence. Training those signals together could produce systems that understand cause and effect more deeply than models trained on any one format.
That matters beyond better clips. The same foundation could eventually power interactive editing, simulations, computer control, factory automation, and robots that learn new tasks with less specialized training.
Our take: The most important FLUX 3 demo may not be its prettiest video. It may be a robot handling a floppy cable on a factory floor. Generative video has quietly become training for physical intelligence, because faking reality convincingly requires learning some of its rules.
BFL still has plenty to prove outside its own evaluations. But if this architecture scales, the line between “media model” and “robot brain” is about to get very blurry.

FROM OUR PARTNERS
Join the OutSystems developer community and start using AI to develop, deploy, and scale your next mission-critical agentic application for free. Go from prompt to production faster, with full control, on a unified, agile, and enterprise-proven platform.

🎓 AI Skill of the Day: Run a 25-Minute Agent Sweep to Protect Your Cognitive Load
Running more AI agents when running your own “software factory” can create a new problem: you spend all day checking whether they finished… and wth they just did.
In Ryan Carson’s conversation with Greg Isenberg, Carson shares a better system: pin the important threads, then review them on a roughly 25-minute cadence. Everything else can wait.
Try it with ChatGPT, Claude, Codex, Devin (Ryan’s pick), or any tool where several tasks can run separately:
Give each agent one clearly defined outcome.
Pin only the tasks that must move today.
Let them work without constant interruption.
Every 25 minutes, request the same short update: progress, blocker, evidence, and next action.
Approve, redirect, or stop the task. Then leave again.
The goal is to keep agents moving while protecting your attention. Carson’s favorite insight: once agents handle execution, your bottleneck becomes judgment rather than typing.
Paste this into each important thread:
Continue working independently until you either finish or need a decision from me.
When I check back, report only:
1. Current status
2. What you completed
3. Evidence that it works
4. Any blocker or decision you need from me
5. Your recommended next action
Do not wait for approval unless the next step is destructive, irreversible, security-sensitive, or changes the agreed scope.Friendly tip: If you’re doing this in Claude Code, turn on Auto Mode in a Cloud Environment instead of running locally. In Codex, use “Approve for me.”
For loads more tips on setting up advanced coding workflows as a solopreneur, read our full article, then check out the video.
Have a specific skill you want to learn? Request it here.

🍪 Treats to Try
Handoff H1 reads construction blueprints and generates material takeoffs across trades, matching experienced estimators on the company’s benchmark — Scale and Enterprise.
Compound Engineering v3.20 routes planning, implementation, and adversarial review across different models while carrying context between agent sessions — free and open tooling.
Roboto Agents traces robot failures from telemetry to the responsible code, tests hypotheses against real runs, and drafts fixes — no public pricing.
Grok Build Workflows breaks large jobs into plans, runs up to 1,024 agents in parallel, and saves workflows as shared slash commands — available through the xAI CLI.
PureBox scans your Gmail and splits it into Attention, Archive, and Trash, so you can bulk-approve the clutter instead of opening 3,000 emails one by one—free trial, then $4.99/mo.
Aymo AI bundles GPT-5, Claude, Gemini, and 50+ other models into one workspace, so you run the same prompt across all of them and skip paying for six subscriptions—free trial, then $4/mo.
Openbase lets you talk to Claude Code or Codex instead of typing, so you say "fix the login bug" from your phone, approve the commands, and review the diff when it's done—no pricing details (waitlist).

New from The Neuron: AI Explained
New episodes air every week on Wednesdays: Spotify | Apple Podcasts | YouTube

📰 Around the Horn
NVIDIA, Microsoft, Meta, and other AI companies urged targeted enforcement against misuse instead of broad restrictions on downloadable model weights.
OpenAI reportedly waited about ten days to tell Hugging Face its models were behind the July 11 intrusion after agents escaped a sandbox.
Anduril was reportedly in talks to raise money at a valuation near $100B.
Amazon began requiring sellers to label ads featuring AI-generated people after a New York synthetic-performer law took effect.
On the Rate Limited pod, Ray Fernando described a personal improvement ecosystem that blends an agent, texting, and a custom app around one person’s goals, while Nathan Snell said users care about outcomes, not the AI label. GOU Coder Adam Larson argued for automating low-impact code aggressively while scrutinizing systems that can affect every customer.

😹 Monday Meme

A Cat’s Commentary


![]() | That’s all for now.
|
P.S: Before you go… have you subscribed to our YouTube Channel? If not, can you?
P.P.S: Love the newsletter, but only want to get it once per week? Don’t unsubscribe—update your preferences here.






