Four years ago, Salesforce's own customer data were scattered across 650 different data streams and 266 million fragmented profiles, so getting a complete customer picture felt impossible.. My guest today, Salesforce’s Chief Data Officer Michael Andrew, fixed that… and now he's rebuilding it all again for a new kind of customer: AI agents. #ad #SalesforcePartner
More on Michael Andrew:
• Has been at Salesforce for nearly 8 years, eventually growing into the CDO role.
• Has spent nearly three decades listening to customers through data and, at Salesforce, he runs one of the largest Data 360 deployments in the world.
• Previously held a range of analytics and data science leadership roles between San Francisco and London.
In this special episode recorded live at Dreamforce two weeks ago, Michael explains:
• Why agents need ten times more data than humans do.
• Why moving data between warehouses is usually money wasted.
• What data scientists, engineers and analysts should be learning right now as agents (rather than people!) become the main consumers of their work.
Filtering by Category: Podcast
Tokenomics: Why Your Agentic AI Bill Is Exploding (and How to Fix It), with Tyler Cox and Ish Shah
Over a single weekend, Ishan S. burned through 2 billion tokens (building a video game for his wife)! He and Tyler Cox join me in today's episode to explain why agentic A.I. bills are exploding... and how a box under your desk can cut them by up to 93%.
Tyler and Ish are returning guests on my podcast... but they are on the show *together* for the first time today. They are both Distinguished Engineers in the Office of the CTO for the client group at Dell Technologies, where they work out how to run powerful A.I. models on the machines closest to you.
In this information-rich episode, Tyler and Ish:
• Dig into "tokenomics" (why agents and their sub-agents devour so many more tokens than chatbots ever did).
• How to pick the right LLM for the job.
• How moving agentic workloads off pay-per-token cloud APIs and onto your own hardware can pay for itself in as little as two months.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce’s Gaurav Pathak
When humans get bad data, we argue about it in meetings. When AI agents get bad data, they can turn it into an answer with total confidence.
My wise guest today, Gaurav Pathak, calls it “garbage in, gospel out.” The takeaway? Better data and context lead to better AI — and, of course, Gaurav has some thoughts on how to get there. #ad #SalesforcePartner
Gaurav Pathak:
• Senior Vice President Product Management AI and Metadata at Salesforce.
• Previously spent 13 years at Informatica building its metadata and AI products, including the CLAIRE AI engine, before Salesforce acquired the company last year.
• Holds a degree in computer science and engineering, as well as an MBA
In this episode, filmed live at Dreamforce in San Francisco last week, Gaurav explains:
• Why context is 95% of the battle for enterprise agents.
• Who the "sin eaters" are that pay for an agent's mistakes.
• The three skills that matter most for AI engineers today.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
How AI Brought a Podcast Back From the Dead, with Linear Digressions’ Katie Malone
Dr. Caitlin Malone's extremely popular "Linear Digressions" podcast went silent for five years. Then A.I. made it possible to bring it back... and to manage it like a team. Hear all about effective agent teams in today's excellent episode.
More on Katie:
• Hosts and produces "Linear Digressions", one of the world's most popular data-science podcasts. (It's excellent — check it out!)
• Senior Researcher at Moonlite AI.
• Previously Sr Director of the A.I. Innovation Lab at the Health Care Service Corporation, Sr Director of Data Science at Tempus Labs, and Director of Data Science at Civis Analytics.
• Has taught machine learning via Udacity and at the University of Chicago.
• Holds a PhD in experimental particle physics from Stanford University focused on working with CERN data.
In this episode, Katie explains:
• Why managing people and managing A.I. agents are the same skill in different clothing.
• Why A.I. could hollow out the expertise we need to catch its mistakes.
• A fascinating range of data paradoxes from Simpson’s to Benford’s.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
The Chip Built for Agentic AI Inference, with SambaNova’s Anton McGonnell
GPUs are the workhorse of A.I. inference… but they aren’t actually optimized for inference! SambaNova has raised over $2B to build a new chip that is, pushing the frontier of real-time A.I. speed and bandwidth. Hear about it from Anton in today's episode.
More on Anton McGonnell:
• VP of Product at SambaNova, a Bay Area A.I.-hardware business recently valued at $11 billion.
• Was previously Director of Product Management for Machine Learning at UiPath and responsible for A.I. research at Glasswing Ventures.
• Holds an MBA from Harvard Business School and a degree in computer science from Queen's University Belfast.
In today's episode, Anton explains:
• Why agentic A.I. has changed the shape of inference workloads.
• How SambaNova's chip sidesteps the memory bottleneck that slows GPUs when generating output tokens.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
OpenAI’s GPT-6 Astra
At GPT-6 Astra's launch, OpenAI president Greg Brockman suggested future observers may point to it as the arrival of AGI. In fact, Artificial Analysis's latest benchmarking suggests it merely catches OpenAI up to Anthropic. Here are the key details:
GPT-6 ASTRA BASICS
• Successor to GPT-5.6 Sol, trained on 100,000+ GPUs at the Stargate facility in Texas.
• It's the first OpenAI model where other A.I. models helped supervise training.
• API pricing is $10/M input tokens and $50/M output (~2.5x Sol's promotional rate; in line with Claude Fable 5.1), with five reasoning-effort settings from low to max.
• Rolled out to ChatGPT Plus, Pro, Business and Enterprise (Enterprise admins must switch it on).
GPT-6 ASTRA CAPABILITIES
• Computer use: 72.6% on OSWorld 2.0 (vs 65.7% for Sol) while finishing tasks ~47% faster.
• Coding: ~58% on Terminal-Bench 4.0 (vs 37% for Sol, ~56% for Fable 5.1), though level with the top Claude models on FrontierCode.
• Math: 99.9% on ARC-AGI-3 (Sol scored under 8%) and 97.6% on FrontierMath Tier 4. Astra also helped tighten a bound on prime gaps from 240 to 186.
• Science: ~65% on Terminal-Bench Science (vs ~22% for Sol, ~53% for Fable 5.1) and ~65% fewer output tokens than Claude Opus 5 on Agents' Last Exam.
• *All of the above stats, however, come from OpenAI themselves.* The third-party view (see chart) shows Astra at 53, tied with Claude Fable 5.1 (and up from Sol's 47) on Artificial Analysis's (widely respected) composite "Intelligence Index".
SAFETY
• First OpenAI model rated "Critical" for cybersecurity under their Preparedness Framework: without safeguards it finds unknown vulnerabilities and builds working exploits (100% on ExploitBench vs 78.5% for Sol; two zero-days found in a fresh Chrome evaluation).
• The release was slowed and gated. The public version supports defensive work but refuses exploit writing; vetted orgs in OpenAI's "Daybreak" program get fewer restrictions.
• New misalignment monitors can pause or halt actions that look unauthorized.
• Encouraging: on a scope-creep evaluation, Sol exceeded its authorized target 48% of the time; Astra did so 0% of the time.
• Concerning: Astra's written reasoning is harder to monitor than Sol's.
IS IT AGI?
• AGI isn't a binary event; there are degrees of breadth and depth (listen to Episode #748 of my podcast for more details).
• BOTTOM LINE: Astra looks like a noteworthy jump in general-intelligence capability for OpenAI relative to their previous flagship model (5.6 Sol), but it does not appear to be moving the frontier toward AGI in any major way... it is merely catching up to Anthropic at the frontier vicinity.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano
Today, the renowned A.I. scientist Dr. Luis Serrano (>200k YouTube subs; ex-Apple; ex-Cohere) returns to launch the second edition of his bestselling ML book and for a mind-bending convo on how LLMs "curve space-time".
More on Luis:
• Founder of the Serrano Academy (YouTube channel with over 200,000 subscribers hooked on his visual, intuitive explanations of machine learning).
• Previously worked at Apple, Cohere and the quantum-computing startup Zapata.
• The second edition of his bestselling Manning Publications Co. book "Grokking Machine Learning" is out this month!
In today's episode, Luis details:
• His mind-bending new paper on how transformer architectures mirror Einstein's curved space-time.
• Why RAG isn't an agent.
• How GRPO (the reinforcement-learning technique behind DeepSeek's reasoning breakthrough last year) works.
• ...and much more!
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
In Case You Missed It in August 2026
If you're in the Northern Hemisphere, summer is drawing to a close... but, ICYMI, today's episode highlights the best bits of the hot hot hot conversations that we had on my podcast in August:
1. In one of the technically richest conversations ever on the show, MongoDB's Field CTO of A.I., Pete Johnson, details the two common traits every company seeing ROI on A.I. investment have.
2. Gurobi Optimization's manager of decision-intelligence strategy Jerry Yurchisin returns to the show to explain where the division of labour between agents and mathematical solvers ought to fall.
3. Priyanka "The Cloud Girl" Vergadia (bestselling author, ex-Google, ex-Microsoft) walks us through how she structures Claude skills so that her output stops being slop.
4. dbt Labs founder and CEO Tristan Handy explains why the semantic layer matters more, not less, now that analytics agents are the ones asking the questions.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Agentic AI Skills That Matter Now, with Aishwarya Srinivasan
My exceptional guest today is world-renowned data scientist, author, tech entrepreneur and content creator (>1.2m followers!) Aishwarya Srinivasan... and she does not disappoint! This is one of my favorite episodes ever, enjoy :)
More on Ash:
• Co-founder of The Gen Academy and a stealth A.I. startup.
• Prolific A.I. startup advisor and investor.
• Sought-after global keynote speaker.
• Author of the book "What's Your Worth? Discovering Your Personal Brand".
• Has held roles at Nebius, Fireworks AI, Microsoft, and been a data scientist at Google, IBM and Goldman Sachs.
• Holds a Master's in Data Science from Columbia University.
In this episode, Ash covers:
• Where defensibility comes from when code is nearly free.
• How to evaluate non-deterministic agentic systems end to end.
• Why reinforcement learning is having such a resurgence.
• How to future-proof your career.
• ...and much, much more!
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents
Ever written careful instructions for an A.I. agent, only to watch it ignore them an hour into a session? The fix lies in knowing WHERE your instructions should live. Read on for all the details...
WHY AGENTS "FORGET"
• Everything an agent knows lives in its context window, and every token costs money and (more subtly) attention.
• Long sessions trigger "compaction": The conversation gets summarized to free up room, and instructions given early can get squeezed out.
• Every steering method answers one question: How do I make an instruction cheap to carry and hard for the agent to lose?
THE STEERING TOOLKIT (drawn from Anthropic's Claude Code, but the ideas generalize)
1.) Always-on files (e.g., CLAUDE.md): Loaded every session and re-read after compaction. Keep them under 200 lines, give them an owner, review changes like code.
2.) Rules: Path-scoped constraints that load only when relevant files are touched.
3.) Skills: Procedures (deploy workflows, checklists) whose full text loads only when invoked.
4.) Subagents: side tasks that run in isolated context windows... only the final summary returns.
5.) Hooks: deterministic code that fires on lifecycle events. The big idea: An instruction is a probability while a hook is a guarantee. If something must *never* happen, enforce it with code, not prose.
THE INDUSTRY IS CONVERGING
• AGENTS.md (kicked off by OpenAI, stewarded by the The Linux Foundation and backed by Google, Microsoft and AWS) serves similar function to CLAUDE.md and is read natively by Codex, Cursor, GitHub Copilot, Gemini CLI and dozens more tools across 60,000+ repositories.
• ETH Zurich researchers studied 138 real-world repos: Developer-written instruction files improved agent task success ~4% and cut agent-introduced bugs by 35-55%.
• The same study found LLM-generated instruction files DECREASED success while raising inference costs by 20%+. These data suggest the value is the human judgment encoded in the file... **you can't delegate the steering wheel to the thing being steered**.
NOT A CODER? THE SAME FRAMEWORK APPLIES
• Custom instructions = your always-on file (keep it short).
• Projects and Gems = your path-scoped rules.
• Custom GPTs = the consumer cousin of skills.
BOTTOM LINE
• Match persistence to relevance: always-on files stay ruthlessly short; procedures and area-specific conventions load on demand.
• "Always" and "never" are signals you need a guardrail (deterministic enforcement), not an instruction.
• Treat steering files as code: owned, reviewed and pruned. An instructions file that grows without gardening dilutes the instructions that matter.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
How dbt Won Analytics Engineering, with dbt Lab’s CEO Tristan Handy
dbt is THE open-source tool that brought software-engineering rigor to data transformation; it's now used by over 100,000 teams. Today's rockstar guest, Tristan Handy, is CEO of dbt Labs, the company behind the movement.
More on Tristan:
• President and co-founder of "Fivetran + dbt Labs" (recent merger).
• Coined the term "analytics engineering".
• Over two decades of experience as a data practitioner working in both large enterprises and startups.
• His expertise and data industry best practices have influenced thousands of subscribers and listeners weekly via his newsletter (The Analytics Engineering Roundup) and The Analytics Engineering Podcast.
In this episode, Tristan explains:
• What dbt is.
• Why dbt Labs' the recent merger with Fivetran is a win for dbt users.
• Why skill files are so powerful.
• Why he turned down acquisition offers for years
• How the semantic layer keeps A.I. agents from confidently getting your metrics wrong
• ...and much more.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
How to Choose Model Size and Effort Level: The Two Critical Dials
Every major A.I. platform now offers two dials that shape what you get back: how large of a model you pick and how much effort it can invest. Here's what each dial does and how to use them together:
THE "MODEL SIZE" DIAL
• Model selection swaps which set of frozen weights handles your request. Weights are fixed at training time; nothing in your prompt changes them.
• Larger models encode more knowledge and capability, and each output token costs more.
• Your prompt steers predictions but doesn't teach: Paste in docs for a library the model has never seen and it will use them for that request, then retain nothing.
THE "EFFORT" DIAL
• Effort shapes all output tokens: reasoning, tool calls and messages to you.
• At high effort, a model reads more files, verifies more of its own work and pushes further before checking in. In one Anthropic comparison, the high-effort path generated ~7x more tokens to reach a higher-confidence answer.
• Effort sets how far a model is willing to travel, not how far it *must* (thus a well-trained model stops when it finds the bug rather than padding the bill).
WHICH DIAL SHOULD YOU TURN?
• First off, when output disappoints, try fixing your prompt and context. A vague request is the most common culprit for a disappointing output and *no* knob will fix that.
• Skipped files, unrun tests, an abandoned refactor? Those are examples of DILIGENCE FAILURE: in these cases, RAISE THE EFFORT.
• Full context, visible attempt, still confidently wrong? Those are examples of CAPABILITY FAILURE: in these cases, upgrade to a BIGGER MODEL.
COST
• Cheaper *per token* isn't always cheaper *per task*. On hard multi-step work, a large model reaches the quality bar in fewer steps, so total cost can come out lower than a small model grinding at the ceiling of its ability (see chart).
• On routine work, the equation flips: Both models get it right, so the big model's extra verification is wasted money. Drop down to a smaller model.
WHAT CAN YOU DO?
• Start with defaults; providers tune them to what most users want to spend. Treat effort as a preference for your kind of work, generally NOT as a task-by-task fiddle.
• Route routine work to smaller, cheaper models and reserve the frontier for problems that stretch it.
• This applies everywhere: Anthropic, OpenAI, Google and open-source reasoning-model providers have all converged on separate capability and effort controls.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Alibaba’s Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs
Chinese tech giant Alibaba is shipping Qwen3.8-Max: It'll be the largest open-weight A.I. model ever released and (as shown in chart) it's competing at the frontier alongside closed-source models from Anthropic and OpenAI. Here's what you need to know:
THE MODEL
• 2.4 trillion parameters in a mixture-of-experts architecture; only ~95B are active per token, so the headline number reflects capacity, not per-request compute.
• Accepts text, images and video, with a one-million-token context window (~750K words).
• Selectable low, medium or extra-high reasoning effort, so no paying for lengthy thinking traces when you need a quick lookup.
• All part of Alibaba's ~$53B, three-year bet on cloud and A.I. infrastructure.
THE BENCHMARKS
• Alibaba frames it as second only to Anthropic's Claude Fable 5; independent signals so far land in a similar neighborhood.
• Immediately became the highest-ranking Chinese model for text on the Arena leaderboard and ranked second globally on vision.
• Scored 86.6 on Terminal-Bench (agentic command-line tasks), ahead of both Claude Opus 4.8 and Fable 5.
• Vendor demos showcase multi-day autonomy: 10+ days coding unattended, plus reproducing an ML research paper from scratch and then beating 87% of 526 human teams in a live contest.
THE PRICE WAR
• $2 per million input tokens, $6 output and 25¢ cached input.
• Undercuts domestic rival Kimi K3 by more than half on output.
• Combined rate is less than a third of Claude Opus and under a quarter of GPT-5.6 Sol's... and cached-input pricing lets agentic workloads collapse toward the 25¢ floor.
IS IT SAFE TO USE A CHINESE MODEL?
• The risk depends less on the model and more on how your data reach it.
• Consumer apps and the hosted API route data through infrastructure governed by Chinese law, so keep anything sensitive or proprietary out of that.
• Safer: open weights hosted by a Western A.I. cloud (like Lightning AI) in your own jurisdiction.
• Safest: run the model on your own hardware: weights are inert files that can't phone home.
• Caveats: outputs reflect training under Chinese content regulations, so evaluate before trusting; practice supply-chain hygiene (official repos, checksums) and check whether your industry restricts Chinese-origin models.
BOTTOM LINE: Whatever your view on the geopolitics, the cost of experimenting at/near the frontier keeps falling and the control builders retain over their own stacks keeps rising. In my view, that's great news :)
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Vector Search, Agentic Memory and Effective RAG, with MongoDB’s Pete Johnson
Today's guest, MongoDB's "field CTO for A.I." Pete Johnson, is exceptional... don't miss this episode! We get deep into vector search, agentic memory, RAG and much more, with Pete vividly explaining technical content like no other.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Mathematical Optimization in the Agentic AI Era, with Gurobi’s Jerry Yurchisin
LLMs will confidently tell you they've optimized your entire business while ignoring the one constraint that could cost you millions. Today's episode with Jerry Yurchisin is all about a technique that makes breaking a constraint mathematically impossible.
More on Jerry:
• Manager of Decision Intelligence Strategy at Gurobi Optimization, a "mathematical optimization" solver that's used by the vast majority of Fortune 100 companies.
• Has over a decade of experience in operations research, data science, and visualization... and specializes in enhancing decision-making.
• Prior to Gurobi, Jerry worked in consulting (OnLocation, Inc. & Booz Allen Hamilton) where he focused on mathematical optimization, machine learning, statistics and simulation.
• Taught statistics and operations research at The University of North Carolina at Chapel Hill and graduate math at Ohio University (he also holds Master's degrees from both of these institutions).
In this episode, Jerry:
• Lays out where mathematical optimization fits in the agentic AI era (hint: LLMs formulate problems and solvers like Gurobi guarantee the answers).
• Shares striking mathematical-optimization applications spanning energy grids, retirement planning... and the model that powered @null's women's team to a gold medal at the Paris Olympics.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
OpenAI Agent Breaches Hugging Face: All You Must Know incl. How to Protect Yourself
Cheesy secret "agent" thumbnail? Oh yes! As we've all surely heard by now, an OpenAI agent escaped its sandbox and hacked Hugging Face's servers. With the dust settled, here's *everything* you need to know, incl. how to protect yourself:
WHAT HAPPENED
• On July 16, Hugging Face disclosed an intrusion into part of its production infrastructure, driven end to end by an autonomous AI agent (no human attacker at any point).
• Five days later, OpenAI revealed the attacker: its own models (GPT-5.6-Sol plus an unreleased, more capable model), running a cyber-capability evaluation with safety guardrails switched off.
• Nobody told the models to attack anyone. Tasked with solving a security benchmark (ExploitGym), they exploited a zero-day in their sandbox's package proxy, escaped to the open internet and went after Hugging Face to steal the benchmark's answers. (Paperclips anyone? 🖇️)
WHY IT MATTERS
• The models found and exploited novel attack paths in real-world production systems without source-code access.
• Hugging Face is a $4.5B firm with a dedicated security team; a less-resourced target would have fared far worse.
• The warning signs were documented months earlier: in ExploitGym testing, frontier agents routinely abandoned the assigned vulnerability and invented their own route to the flag.
THE TWIST
• Hugging Face detected and dissected the breach using AI... LLM-driven forensics compressed days of analysis into hours.
• But the commercial APIs from American frontier labs they tried first (e.g., OpenAI's or Anthropic's) refused to process the attack logs. Safety guardrails "cannot distinguish an incident responder from an attacker", they said.
• The forensics ran instead on a Chinese open-weight model (Z.ai's GLM 5.2) on Hugging Face's own infrastructure, which also kept attacker data in-house.
LESSONS 1 (IF YOU BUILD AGENTS)
• Treat your egress allowlist as attack surface, not a safety guarantee. Pre-install dependencies and air-gap where you can.
• If you disable safeguards for testing, scale containment and monitoring in proportion.
• Log every agent action and alert on outbound connections that shouldn't exist.
LESSONS 2 (IF YOU DEFEND A PLATFORM)
• Rotate your Hugging Face access tokens and review recent account activity today.
• Treat user-supplied data and model artifacts as executable content, not inert files. Audit every code-execution path in your pipeline.
• Stand up a capable open-weight model on your own infrastructure and validate it for forensic log analysis BEFORE you need it at 2am (this is easy to do with, say, Lightning AI).
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Weapons of Math Destruction, Ten Years On, with Dr. Cathy O’Neil
What makes an algorithm terrifying? My guest today, mega-bestselling author of "Weapons of Math Destruction" Dr. Cathy O'Neil, says it's not the complexity of the math; it's the secrecy, the unaccountability and the fact that you can't opt out.
More on Dr. O'Neil:
• A decade after her "Weapons of Math Destruction" (2016) sounded the alarm on algorithmic harm, she's busier than ever.
• Through her algorithmic-auditing firm ORCAA and her nonprofit OCEAN, she now provides the statistical evidence behind lawsuits against some of the world's biggest tech companies.
• Co-hosts the "A.I. Skeptics" podcast.
• Also wrote "The Shame Machine: Who Profits in the New Age of Humiliation", which was published in 2022 (like WMD, also by Penguin Random House).
• Before writing trade publications, her first book was actually an O'Reilly book, "Doing Data Science".
• Earlier in her career, she held academic positions at Massachusetts Institute of Technology and Barnard College before becoming a Wall-Street analyst at The D. E. Shaw Group.
In this episode, Cathy:
• Punctures A.I. hype.
• Explains why A.I. won't so much replace workers as degrade them.
• Lays out how all of us can demand accountability.
I wanted to have this exceptional conversation with Dr. O'Neil for a decade, enjoy!
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
The Open-Weight 2.8-Trillion Parameter Competing at the Frontier
A new model out of China, Kimi K3, has rattled investors globally, kicked off a pricing skirmish among the big American A.I. labs and reignited the open-weight A.I. debate. Here's everything you need to know:
(The attached chart from The Economist tells the story at a glance: Open-weight models like K3 are now nipping at the heels of the closed-weight frontier.)
THE COMPANY BEHIND IT
• Moonshot AI is a Beijing-based startup backed by Alibaba, with a $20B+ valuation and reported annual recurring revenue above $200m.
• K3 is something of a comeback: Moonshot's market position had eroded following DeepSeek's rise last year... now the student of that disruption has become the disruptor.
WHAT IS KIMI K3?
• A 2.8-trillion-parameter model that Moonshot claims is the largest open-weight A.I. model in the world (it is also the largest I'm aware of).
• It's a "mixture-of-experts" architecture: only 16 of 896 "expert" submodules activate per token, so inference costs are far lower than the headline parameter count suggests.
• Features a one-million-token context window, native visual understanding and two architectural innovations ("Kimi Delta Attention" and "Attention Residuals") that reportedly deliver ~2.5x better scaling efficiency vs. the K2 generation.
HOW GOOD IS IT?
• Moonshot itself says K3 trails Claude Fable 5 and GPT-5.6 Sol overall, but beats the next tier down (Claude Opus 4.8, GPT-5.5) on coding and agentic benchmarks (see chart).
• Independent signals are encouraging: K3 scores 57 on the Artificial Analysis Intelligence Index (median for its price tier: 31) and topped Arena's front-end coding leaderboard.
THE PRICING SHAKE-UP
• K3 costs $3 per million input tokens and $15 per million output tokens, undercutting Claude Opus 4.8 ($5/$25) and GPT-5.6 Sol ($5/$30).
• Cache-hit input tokens cost merely 30¢ per million, which is huge for agentic and RAG workflows.
• OpenAI and Anthropic have already responded by expanding token allowances to retain users.
WHAT CAN YOU DO?
• If model weights ship under the promised Modified MIT license (this is expected next week), any will be able to run very-near-frontier-class A.I. on their own infrastructure with no per-token fees and no data leaving their walls.
• Every price war between labs is a subsidy for the applications you're building... the cost of experimenting with world-class A.I. has never been lower and strong open-weight releases like this will continue to bring price pressure in your favor 😎
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
The Math Still Matters: Deep Skills in the Age of AI, with Dr. Catherine Williams
Dr. Catherine Williams was solving black-hole equations with pen and paper before she ever wrote a line of code and, in today’s episode, she makes the case that going deep on AI/ML math matters more than ever...
...even now that AI can do the math for you.
More on Dr. Williams:
• PhD in math researching general relativity and black holes.
• Postdocs at Stanford University and Columbia University.
• Became one of the very first data scientists when she joined AppNexus back in 2012, around the same time "data scientist" became a job title.
• Across more than a decade of senior data leadership at AppNexus, Xander, Qualtrics and now Candid, she's watched our field get born and then reinvent itself again and again.
In today's episode, Catherine traces the data science and AI evolution — from Bayesian models to BERT to today's LLMs — and shares sharp guidance on which skills will still matter as machines take over more of the technical work.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.
Fable 5 as Advisor: Anthropic’s Two-Model Pattern for Smarter, Cheaper Agents
Want near-frontier A.I. agent quality at a fraction of the cost? Anthropic recently productized the Advisor Strategy that pairs a cheap "executor" model with a brilliant "advisor" to give you the best of both worlds:
HOW IT WORKS
• A fast, cheap model (e.g., Claude Haiku or Sonnet) runs the entire agent loop: calling tools, writing code, drafting output.
• A frontier model (e.g., Claude Opus or Fable) sits on standby as a "tool" the executor can consult (like a junior worker phoning their supervisor when unsure).
• Everything happens inside one API call: Anthropic's servers hand the advisor the full conversation transcript and return just 400-700 tokens of advice, making this fast and inexpensive (it's also usually only a one-line code change so it's easy to implement).
THE RESULTS
• Sonnet + Opus advisor beat Sonnet alone on the "SWE-bench Multilingual" benchmark by 2.7 percentage points while cutting cost per task by 11.9%. Better quality AND slightly lower cost.
• Unsurprisingly, the biggest gains come from pairing a very fast/cheap model with a much more capable advisor: For example, on BrowseComp (web research benchmark), Haiku alone scored 19.7%; Haiku + Opus advisor scored 41.2% (more than double!) at 85% less cost than Sonnet alone.
• Newest data, from last week: On "SWE-bench Pro", Sonnet 5 + a Fable 5 advisor captured ~92% of Fable's standalone performance at ~63% of its cost.
WHY IT WORKS
• The advisor's output is tiny relative to the whole task, and a good plan delivered early prevents wasted attempts and misguided tool calls.
• Unlike OpenAI's router (which dispatches queries to a model up front), the cheap model runs the show and escalates itself mid-task with full shared context.
PRACTICAL LESSONS
• Skip it for single-turn Q&A; it shines on long-horizon agentic work (like coding, research, computer use).
• Executors under-call the advisor by default so prompt them to consult it early (before committing to an approach) and late (before declaring the task done).
• Cap advisor output at ~2,000 tokens (~7x cost reduction, no quality loss) and enable prompt caching for long loops.
• The pattern is spreading: OpenRouter now offers a cross-provider version (e.g., a Google Gemini executor consulting Claude).
• Alternative design patterns such as having a powerful "orchestrator" (shown below the advisor pattern in the chart I included in this post) might work even more effectively for your use case so it could be worth comparing them.
BOTTOM LINE
Frontier A.I. progress is no longer just bigger models... it's smarter economics in composing the models we already have.
The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.