Ms. Wolf logo

Ms. Wolf

Archives
Log in
Subscribe
September 22, 2026

New Episode Ready: AI & Marketing Research Radar — 2026-09-22

New Episode Ready

AI & Marketing Research Radar

2026-09-22  ·  AI and marketing  ·  12 papers screened  ·  3 selected

▶  Listen to This Episode

Apple Podcasts  ·  Spotify  ·  Buzzsprout


First-pass research briefing, not a final academic review. Always read the original paper before citing.

Paper A

Levels of AGI for Operationalizing Progress on the Path to AGI

Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin et al. — 2024 — International Conference on Machine Learning (ICML 2024)

conference paper  ·   ·  skip

https://arxiv.org/abs/2311.02462

Key findings

  • The paper proposes a tiered system for ranking AI — think of it like a video game level system — where AI moves from 'Emerging' to 'Competent' to 'Expert' to 'Virtuoso' to 'Superhuman,' rated on both how good it is at tasks (depth) and how many different kinds of tasks it can handle (breadth).
  • Current AI tools like ChatGPT sit somewhere around 'Emerging to Competent' on this scale — they can do many things fairly well, but they are not yet close to reliably outperforming humans across all domains.
  • How much control a human keeps over an AI system (called 'autonomy level') is treated as a separate, adjustable dial — not something fixed by the AI's capability level. The same powerful AI can be deployed with heavy human oversight or with near-total independence, and that choice matters enormously for safety.
  • The Turing Test and other famous AI benchmarks are rejected as insufficient ways to measure AGI, because they test whether an AI sounds human rather than whether it can actually accomplish useful tasks reliably.

Marketing implications

  • When evaluating AI tools for your marketing team, use the paper's two-part test as a mental checklist: How well does this tool perform at the specific task you need (depth)? And how many different tasks can it handle without breaking (breadth)? A tool that scores high on both is genuinely more capable — not just better-marketed.
  • The paper's point about autonomy being a separate choice from capability is directly useful: just because an AI can do something doesn't mean you should let it run unsupervised. Set explicit human-review checkpoints in your AI workflows — especially for anything customer-facing — regardless of how capable the tool seems.
  • Use this framework to cut through vendor hype. If a vendor claims their tool is 'AGI-level,' ask which level on what tasks. Vague claims about 'human-level AI' are not meaningful — what matters is whether it reliably outperforms alternatives on your specific use case.

Paper B

On the Measure of Intelligence

Francois Chollet — 2019 — arXiv (preprint)

preprint  ·   ·  skip

https://arxiv.org/abs/1911.01547

Key findings

  • Current AI systems are judged by how well they do on specific tasks (like chess or video games), but this doesn't actually measure intelligence — it just measures how much training data and built-in knowledge they were given. You can make any AI look smart by giving it enough practice on a specific task, which hides whether it can actually learn new things.
  • The author argues real intelligence is about how quickly and efficiently a system can learn a new skill from very little experience — not how good it already is at tasks it was trained on. Think of it like judging a student by how fast they pick up a brand-new subject, not by how many subjects they've already memorized.
  • The paper proposes the Abstraction and Reasoning Corpus (ARC): a set of puzzle-like tests that use only the kinds of basic knowledge all humans are born with (like understanding objects, counting, and spatial relationships). The idea is that ARC can test whether an AI can reason like a human without having been pre-trained specifically on those puzzles.
  • Most AI today, including large models, cannot reliably solve ARC tasks that an average human can solve in seconds — which suggests current AI systems, despite impressive-sounding benchmarks, may be missing a core ingredient of human-like intelligence: the ability to generalize from very few examples.

Marketing implications

  • When a vendor claims their AI is 'as smart as a human' because it scored well on some benchmark, ask what the benchmark actually tests. This paper explains why task-specific scores can be gamed and don't tell you much about whether the AI will handle something new — which is exactly what marketing teams face constantly.
  • If you're evaluating AI tools for your marketing stack, test them on tasks they haven't been specifically designed for. An AI that can only do what it was trained to do will break the moment your campaign needs something slightly different. Ask vendors: 'What happens when we use this outside its training domain?'
  • Be skeptical of AI capability claims that rely on leaderboard rankings or benchmark scores without context. The paper's core argument — that training data volume can 'buy' high scores without real intelligence — is directly relevant when assessing whether an AI tool will be flexible enough for creative or strategic marketing work.

Paper C

Measuring Progress Toward AGI: A Cognitive Framework

Ryan Burnell, Yumeya Yamamori, Orhan Firat, Kate Olszewska et al. — 2026 — arXiv (preprint) - Google DeepMind

preprint  ·   ·  skip

https://arxiv.org/abs/2605.28405

Key findings

  • There is currently no agreed-upon scientific framework for measuring how close AI systems are to human-level general intelligence, which leads to both hype and underestimation of AI capabilities.
  • The authors propose breaking 'general intelligence' into 10 specific mental abilities: perception, generation, attention, learning, memory, reasoning, metacognition, executive functions, problem solving, and social cognition — each of which can be tested separately.
  • To fairly compare AI to humans, the authors say tests must be private (so AI can't have memorized the answers), independently verified, and varied in difficulty and format — similar to how a fair school exam would be designed.
  • The framework produces a 'cognitive profile' for each AI system, similar to a report card, showing where the AI is strong (e.g., reasoning) and where it still falls short of humans (e.g., social cognition or metacognition).

Marketing implications

  • When an AI vendor claims their tool is 'general purpose' or 'approaching AGI,' use this framework as a checklist: ask which of the 10 cognitive areas (e.g., social cognition, metacognition, memory) the tool actually performs well on versus where it has gaps — don't accept vague capability claims at face value.
  • If your team is evaluating AI tools for complex marketing tasks (like writing strategy, managing campaigns, or interpreting customer behavior), this 10-faculty breakdown gives you a practical vocabulary for scoping what the AI can and can't reliably do.
  • For agencies pitching AI-powered services, being able to articulate AI capability gaps — using a framework from Google DeepMind researchers — can help set realistic client expectations and avoid overpromising.

▶  Listen to This Episode

Apple Podcasts  ·  Spotify  ·  Buzzsprout

AI & Marketing Research Radar — Big Plans Media — 2026-09-22

Don't miss what's next. Subscribe to Ms. Wolf:
← Newer New Episode Ready: AI & Marketing Research Radar — 2026-09-22 Older → New Episode Ready: AI & Marketing Research Radar — 2026-09-18
Powered by Buttondown, the easiest way to start and grow your newsletter.