- The Neuron
- Posts
- 😺 3,000 Mexican exam scores wiped over AI
😺 3,000 Mexican exam scores wiped over AI
PLUS: Alibaba's new AI codes alone for 10 days straight.

Welcome, humans.
Alibaba just dropped Qwen3.8-Max, a 2.4 trillion parameter model with a coding agent that reportedly works unsupervised for 10+ days straight, from empty folder to finished product. Next week it's going open-source, alongside a smaller sibling, so you'll be able to run it yourself instead of just reading about it.
The wildest claim in the announcement: it simulated 365 days of e-commerce strategy in one shot, basically living an entire year as a digital shopkeeper before you've finished your coffee. Somewhere, a supply chain manager just had a small existential crisis.
Here’s what happened in AI today:
😺 Mexico's top university, UNAM, canceled thousands of exam scores after results suggested widespread AI-assisted cheating on its first-ever online entrance exam.
📰 Gemini Spark now handles logged-in web errands inside Chrome.
📰 Google’s security agents helped fix 1,072 Chrome bugs.
🍪 Dreamina now generates controllable videos up to three minutes.
😹 Internet meets an alleged two-meter caregiving spider robot.

😺 Mexico's Top University Cancelled Thousands of Exam Scores Over Suspected AI Cheating
UNAM, Mexico's most prestigious university, just canceled thousands of exam scores over suspected AI cheating, a scandal that has spread all the way to the country's president.
UNAM put its entrance exam online for the first time this year, hoping to make it easier for applicants outside Mexico City. More than 158,000 people took it in May and June. The scores that came back didn't add up.
Here's what happened:
Between 2021 and 2025, about 3.5% of applicants historically scored 100 or more correct answers.
This year, that rate spiked high enough to trigger a formal review.
UNAM annulled roughly 3,000 of the ~150,000 tests over suspiciously perfect scores.
The university suspended new-student enrollment (the semester starts August 10) while it investigates.
Officials suspect a mix of AI tools like ChatGPT and old-fashioned tricks, like leaked questions.
The exam software, built by Mexican company Territorium Life, was supposed to catch exactly this: it blocked test-takers from opening extra browser tabs and used AI to flag suspicious behavior. Students say the system leaned too hard on AI and not enough on humans watching. Territorium Life maintains its software worked as intended, and argues stopping cheating also requires the university to enforce honesty on its end.
Why this matters: High-stakes testing is moving online everywhere, and AI is increasingly the referee. When that referee can't reliably tell a cheater from a hard studier, everyone pays for it, including the students who did nothing wrong. President Claudia Sheinbaum, a UNAM graduate herself, has now weighed in publicly, turning a testing glitch into a national story.
Our take: Nobody built a system that can tell the difference between a cheater and someone who just studied hard, and that gap is the actual scandal here. Turns out "trust but verify" gets messy fast when the verifier is an algorithm nobody fully understands.

FROM OUR PARTNERS
The integrated coworker for AI native teams
Empower your team to do their best work with Adapt, the integrated coworker that works alongside your team in Slack and deeply understands your business.
Here’s how Adapt is different
Set up takes minutes: connect your tools, add to Slack, and it’s right there for anyone to tag @Adapt for help
Does real, high-ROI work: automates work on a schedule; builds internal tools with live data; and does complex, multi-tool tasks on demand
Learns your business as you work, becomes your company brain
Uses the best AI model for the task, not tied to a single provider
SOC2 Type II, RBAC, and support for personal and company-wide integrations

🎓 AI Skill of the Day: Build Your First AI Agent Without Coding
Most people build agents backward: they connect a pile of tools, then hope the AI figures out what its job is. Start with an agent brief instead.
An agent is AI that works toward a goal using context, tools, instructions, and approval rules. We unpacked those terms in our two-hour Agents 101 walkthrough and timestamped companion guide: chatbots answer, automations follow recipes, and agents work toward a goal.
Pick one repetitive task with a visible finish line, such as preparing a weekly inbox summary. Then define:
What starts the workflow.
What information it can access.
The steps it should follow.
Which actions require your approval.
How it proves the job is complete.
Start in draft-only mode. It can research, organize, and prepare work, but it cannot send, delete, publish, purchase, or change records without permission.
Test it with one normal example, one missing-information example, and one strange edge case. Quality assurance first; office keys later.
Help me design a safe, reusable AI agent for this recurring task:
[TASK]
Do not perform the task yet. First create an “Agent Brief” containing:
1. Goal: The exact outcome it must produce.
2. Trigger: What starts the workflow.
3. Inputs: The files, messages, tools, and context it may use.
4. Steps: The exact sequence it should follow.
5. Permissions:
- May do automatically:
- Must ask before:
- Must never do:
6. Stop rules: When it should pause, ask a question, or return control to me.
7. Success check: The evidence proving the task is complete and correct.
8. Output: The required format and destination.
9. Tests:
- One normal example
- One example with missing information
- One unusual edge case
Default to draft-only mode.
Do not send, delete, publish, purchase, contact anyone, or change external records without my explicit approval. Use the minimum access needed, and clearly flag assumptions, missing information, and uncertainty.
After I approve the Agent Brief, walk me through setting it up in [ChatGPT / Claude / OTHER TOOL] without requiring code.Favorite insight: Your first agent should be boring enough that you can tell when it screws up.
Have a specific skill you want to learn? Request it here.

🍪 Treats to Try
Dreamina creates 30-second AI videos and long-form clips up to three minutes with timestamp controls and up to 50 references (pricing varies by region).
Palette combines video generation, editing, and storyboarding on one multimodal canvas while routing across leading models (credits start at $0.01 each).
Superlinear teaches four practical agent-engineering habits through a free video and podcast series (free to watch or listen).
Cloudflare Kumo gives you accessible interface components with keyboard navigation, focus handling, ARIA support, and Figma token sync (free and open source).
Codex Router runs multiple models side by side inside Codex without replacing official integrations (free and open source).

📰 Around the Horn
Gemini Spark began handling logged-in errands inside Chrome while returning payments and sensitive steps to the user.
Google said its AI security pipeline helped fix 1,072 Chrome bugs and found a 13-year-old sandbox escape.
A survey of managers found 59% used AI for layoffs, and 43% sometimes let it decide without supervision.
Reddit’s lawsuit against Perplexity survived most of the company’s attempt to dismiss the case.
Anthropic revealed that Claude models reached real companies during cybersecurity tests after a configuration mistake left a supposedly sealed environment open to the internet.
A new AI assistant called Orchid was pitched as a fix for forgetful boyfriends who miss anniversaries and let groceries rot, and critics say it risks turning weaponized incompetence into a software category.
General Motors planned a vehicle-native assistant using telemetry, OnStar data, maintenance alerts, and family controls.
Snapchat stopped promoting or rewarding fully AI-generated Spotlight videos.
Want absolutely EVERYTHING that happened in AI this week? Click here!

FROM OUR PARTNERS
Build durable, type-safe AI agents in TypeScript that live right in your codebase.
From streaming LLM responses to custom build extensions, Trigger.dev provides complete runtime control without serverless timeouts or ops overhead. Write your logic, validate, and deploy reliable agents that scale on demand.

😹 Monday Meme

AI can build a working prototype in an afternoon. Apple still wants certificates, privacy answers, screenshots, TestFlight, and a review submission that does not accidentally summon six new error messages.
Corey and Grant walk through the full path from AI-built app to App Store listing, including the confusing steps, mistakes, and delays they hit along the way. Watch the follow-along walkthrough here.

A Cat’s Commentary

This translates to “Everything was okay” in Romanian. And since Romania (statistically speaking) has the lowest life satisfaction rate in the EU, that’s basically a standing ovation from Europe’s toughest crowd!

![]() | That’s all for now.
|
Love robots? We just launched a robotics newsletter! Sign up for it here.
P.S: Before you go… have you subscribed to our YouTube Channel? If not, can you?
P.P.S: Love the newsletter, but only want to get it once per week? Don’t unsubscribe—update your preferences here.









