← All Field Notes
AI September 2026

What 19 sessions at the world’s biggest AI summit actually said

I spent two days at AUTONOMOUS, the Board of Innovation’s virtual summit for people putting AI to work inside real companies. Here’s what I heard from stage, a few things a handful of pharma operators are already doing that I found genuinely useful, and what I think it means for how I build and price my own work.

The short version

  1. Nobody trusts the output.

    AI generates content faster than anyone can check it. A controlled trial of experienced developers found they were 19% slower with AI assistance and believed they were 20% faster. The new bottleneck is verification, and the constraint is human review capacity, not generation speed.

  2. The model is not the moat.

    Everyone has access to the same frontier models, and they’re commoditizing fast. The differentiator is proprietary context — the company history, domain knowledge, and brand truth that gets fed in. Without it, the most expensive model in the world still behaves like a new PhD on day one.

  3. Ideas aren’t the problem. Ownership is.

    When Zalando surfaced 102 candidate use cases in one month, the bottleneck wasn’t creativity; it was deciding who owns what survives. In most organizations, the same person who builds an AI tool maintains it, the exact structure the software industry abandoned in 1968. Without an ownership model, the demo graveyard piles up while nothing reaches production.

  4. AI is repricing the work.

    A personal-injury demand letter costs $2,500 when a junior lawyer drafts it and about $500 when AI drafts it with a lawyer’s review. Per-hour and per-seat pricing assumes that cheaper production means clients pay proportionally less. It doesn’t, it just means clients expect to. The math breaks long before the work does.

  5. The job shifts from making to judging.

    Coca-Cola reports production time in creative work has dropped from roughly 70% of the job to 30–40%, with craft and ideation filling the gap. IKEA compressed a strategic exploration cycle from 18 months to 18 minutes. The value is moving from production to the judgment wrapped around it.

Hit closest to home

Three operators in regulated industries are already doing this

I spend most of my working life inside pharma and healthcare, so three sessions from operators inside Takeda, Merck, and Sanofi landed harder than the rest. Everything else here is general AI practice. This is the part that maps directly onto work I actually do.

Takeda · Saurav Gupta
Head of Global Digital Patient Services

A working demo is not a product

Gupta built an AI tool that wowed leadership in the demo — and got shut down 18 months later because it never moved a business metric. His line: demos run on charisma, products run on certainty. His 5Ps model treats scaling as a funnel where use cases are supposed to fall out at every stage, not a ladder where everything climbs.

Merck · Mariana Hebborn
Global Head of HR AI & Data Governance

Legacy data is an advantage

This runs against the usual “legacy companies are slow” story. Merck has more than 40 years of employee data, giving it a view of talent and skills that younger, data-native companies simply don’t have. The regulated environment is part of what makes that data trustworthy enough to use.

Sanofi · Nitesh Soni
Global AI Solutions Leader, Commercial

Engagement is going digital-first

Medical interactions are moving toward shorter digital touchpoints. HCP engagement now brings scientific and commercial data together in real time, while real-world evidence feeds back into research priorities. His advice on adoption was simple: start with people’s pain points, not the technology.

What I took from this

These aren’t distant examples from companies with different rules than mine. Takeda, Merck, and Sanofi work under the same regulatory pressure I do, and each is already working through what happens after the demo. That’s why the lessons transfer. Gupta is based in Boston, so a peer conversation with someone doing this work inside Takeda feels worth pursuing.

The full findings

Five things people kept saying

Across 19 sessions, five patterns kept surfacing — usually from speakers who’d never met, in very different industries. When that happens, it’s usually a signal worth paying attention to.

1
Verification

Thinking is free. Trusting is not.

AI output is cheap. Verification is the new constraint.

Evidence & implications

Bosch rolled generative tooling out to 12,000 engineers, who self-reported feeling about 20% more productive. Then Bosch’s Head of AI Governance, Jochen Kokemueller, pointed to a controlled trial of experienced developers that found the opposite: they were 19% slower with AI assistance, and believed they were 20% faster. A 40-point gap between feeling and measuring. Code now ships in an hour and then waits three days for review, and Kokemueller’s read is that the constraint on AI adoption is no longer how fast machines generate, it’s how fast humans can verify. His fix borrows from algorithmic trading: machines checking machines, with a named human accountable at the end.

Meta’s Rohit Patel, Director at Meta Superintelligence Labs, described the engineering version: agent output is nondeterministic, so you need evals, a judge model scoring real outputs against curated golden answers, before you can tell whether your agent is getting better or worse.

−19%measured speed with AI, controlled developer trial
+20%self-reported speed, same developers
82:1machine vs. human identities in enterprise, per Cyera
From Meta Superintelligence Labs
01
Build the eval data

Collect 500–1,000 real task examples with actual responses. Build “golden” reference answers through multiple review passes, set a higher bar than any single human answer. Optionally write task-specific rubrics for partial credit.

02
Build the LLM judge

Prompt a judge model with the question, the golden answer, and the response being scored. Optionally fine-tune the judge for your domain. The judge becomes the quantitative compass for everything you ship.

03
Roll judge output into a metric

A single trackable score, as simple as share correct, or a weighted rubric for complex tasks. This is what you watch over time to know whether your agent is improving or drifting.

What this is. The three-step pipeline Meta uses to put a number on agent quality — because without one, every release is a vibe check. Source: Meta AI research; Rohit Patel’s talk.

“Thinking is free. Trusting is not.”

Jochen Kokemueller · Head of AI Governance, Bosch
My take

Pharma taught me verification before AI made it fashionable. Nothing ships without a claim being checked against the source. The interesting question isn’t whether to bolt an AI verification step onto a workflow; it’s whether the domain knowledge behind that check is what actually makes it valuable. That’s the moat. Not the model.

2
Context

Your AI needs context, not more training.

Shared models make current, proprietary context the moat.

Evidence & implications

That line came from Ferrovial’s Dimitris Bountolos and got the loudest reaction of the two days. His argument: models are commoditizing, so the thing that differentiates one company’s AI from another’s is the context around it — company history, domain knowledge, brand truth, kept current. An unconfigured agent is a new PhD on day one: technically brilliant, with zero knowledge of how your company works. Ferrovial’s answer was a federated context platform, deliberately not centralizing the data, but standardizing the meaning attached to it across departments.

And from a different angle, Profound’s Nick Lafferty shared a number with direct consequences for content teams: half of what AI answer engines cite is less than 13 weeks old. Freshness isn’t a nice-to-have; it’s a ranking signal.

“AI does not eliminate ambiguity. It scales it.”

Dimitris Bountolos · Chief Information & Innovation Officer, Ferrovial
My take

This is exactly the wedge I’ve been building Precision AEO around. Content now decays on a roughly 13-week clock in AI search, which makes freshness a recurring job, not a one-time audit. The question for any brand isn’t whether they have content; it’s whether their context is fresh enough to get cited. That’s the whole AEO conversation in one sentence.

Sources & further reading Profound · Ferrovial
3
Ownership

Ideas aren’t scarce. Ownership is.

The bottleneck is ownership, not a shortage of ideas.

Evidence & implications

Zalando’s Natalia Andrievskaya opened planning for a handful of automation workflows. Internal hackathons surfaced 102 candidate use cases in one month. Her answer was a feasibility/suitability scorecard and a triage system, with a named owner on everything that moved forward.

Zendesk’s Mirza Beširović showed the cost of skipping that discipline: one pilot ran $1,500/month in testing and then $1M+/month in production, because nobody modeled real costs before scaling. Melissa Reeve named the underlying structural issue: in most organizations, the same person who builds an AI tool is also the one who maintains it — the exact arrangement the software industry abandoned in 1968.

“Not everyone should be a builder.”

Melissa Reeve · Founder, Hyperadaptive Solutions
Zalando’s prioritization matrix
High AI-Suitability(Should we?) Low
Design to FeasibleValuable, but blockers must be resolved first
Build NowStrong business case and realistic path to automation
Drop or DeferNot attractive enough to pursue now
ParkEasy to automate but lower business value
Low AI-Feasibility (Could we?) High

What this is. Zalando’s two-question filter for 102 candidate use cases. The Build Now quadrant (top-right, highlighted) is where the next sprint goes. Source: Natalia Andrievskaya’s talk; Zalando Tech.

My take

Every team I’ve seen hit this wall the same way. Once people see what these tools can do, the idea flood shows up fast, and it looks like momentum right up until nothing ships. A lightweight version of Zalando’s scorecard, run before a sprint starts rather than after, is cheap insurance against a graveyard full of demos.

Sources & further reading Beat framework · Zalando Tech
4
Pricing

Cheap production breaks hourly pricing.

AI cuts costs faster than hours-based pricing can adapt.

Evidence & implications

BOI’s Amir Ouki put the services-firm version of this on the table directly. A firm that delivers twice as fast can still end up smaller, if clients expect to pay half. His example: personal-injury demand letters, $2,000–$3,000 when a junior lawyer drafts them and about $500 when AI drafts and a lawyer reviews. His prescribed fix is pricing the outcome instead of the hours — Sierra AI now charges per resolved support issue rather than per agent seat.

BOI’s Stefano Orani set the market context: enterprise software spending reached $1.47 trillion, with AI capabilities accounting for roughly 85% of the year-over-year growth. McKinsey’s latest State of AI survey puts two numbers next to each other: 80% of people say AI makes them more productive, but only 6% of organizations tie AI activity to measurable profit. That gap is the story.

6%of orgs tie AI activity to measurable profit (McKinsey)
80%of people say AI makes them more productive (McKinsey)
My take

Agencies and consultancies sell hours and deliverables, and some of those deliverables are exactly the kind AI accelerates. The smart move is repricing from a position of strength, testing an outcome- or scope-based price on one AI-accelerated deliverable you already sell, before a client asks why it still costs the old number.

Sources & further reading Sierra AI · McKinsey State of AI · Gartner top tech trends
5
The work itself

The job shifts from making to judging.

As production shrinks, clients pay for human judgment.

Evidence & implications

Two speakers, different industries, the same shape of change. Coca-Cola’s Dominik Heinrich said production time in creative roles has dropped from ~70% of the job to 30–40%, with craft and ideation time rising to fill the gap. Some companies are rehiring after discovering AI output still needs human refinement to meet the bar.

IKEA’s Explore team compressed a strategic exploration cycle from 18 months to 18 weeks to 18 minutes for a full simulated run — not by working faster, but by killing the decision paralysis that used to force them to choose one path before testing any. BOI’s Laura Stevens added the caveat: organizations default to familiar, low-uncertainty choices, and unless incentives and approval processes change, AI just speeds up the existing bottleneck.

“Eighteen months. Eighteen weeks. Eighteen minutes.”

IKEA’s Explore team · the same exploration cycle at three stages of adoption
My take

This is the framing I keep coming back to. “AI makes us faster” invites the obvious fear. “AI moves our time from production to judgment” is both truer and a better pitch. The value in this work was never the production. It’s the judgment wrapped around it.

Sources & further reading IKEA Innovation · Board of Innovation
If you want to try one

Three places to start

If you read this far and one of these stuck, here are three places to pick up. Each one starts with something you probably already have.

01

Define the verification step before you skip it

Pick the AI-assisted work closest to shipping and write down exactly how the output gets checked before it goes out. Automated where it can be, human where it has to be.

02

Audit your own AI-search freshness

Pick one site you own or write for. Audit how its content currently performs in AI answer engines. Flag where freshness or context gaps are costing you citations and conversations.

03

Reprice one AI-accelerated deliverable

Choose one thing you sell or produce that AI genuinely speeds up. Test an outcome- or scope-based price against the hours-based default. See what conversations change.

Method & further reading

A few notes on method, and links to go deeper

How I put this together. I took notes throughout both days and used AI transcription as a backup, then went back through everything and pulled out the ideas that kept coming up. The summaries here are in my words; the pull quotes stay as close as possible to what was said on stage, and I rebuilt the diagrams from the presented slides. AI was used to assist in organizing and summarizing the content, but the final synthesis and interpretation are my own.

Further reading
About the Author

Jack DeManche

Digital strategy and innovation leader with 15+ years building digital practices for brands including P&G, Microsoft, and Olay. Founder of The Genuine Organization, a strategic services consultancy and creative studio. Builder of Everyday AI and 20+ shipped web applications.