Skip to main content

The World Model Race has an Enterprise Blind Spot

The World Model Race has an Enterprise Blind Spot

About the Author

Steve Ambrose

CCO, Signal Labs

Billions are being invested into AI that predicts how the physical world moves. In this article, we share how the bigger impact will be delivered across institutional world models that predicts how a company moves.


Yann LeCun left Meta in late 2025 after twelve years running its AI research, because he and Mark Zuckerberg disagreed about where intelligence comes from. While Meta keeps scaling language models, LeCun believes that predicting the next word isn't the same as predicting the next state of the world; and only a model which can simulate consequences and plan against them gets to human-level intelligence. Investors put $1.03 billion behind that belief in March, at a $3.5 billion valuation, before his new company had shipped anything.

He's right about the distinction, though Signal Labs is looking at a different "world."

Signal Labs

Receive our latest insights

Three bets, one world

Fei-Fei Li's World Labs shipped Marble last November, which turns a text prompt or a few photos into a 3D scene a designer can walk through. Their newest model called Atlas followed a month ago, able to rebuild a factory floor from a handful of images and simulate the conveyor belts moving. On September 28, AMD agreed to buy the company for about $8.2 billion in stock, seven months after backing a $1 billion round. Then there's Google DeepMind's Genie 3, which creates playable worlds at 24 frames per second so robots and agents can practice before they touch anything physical. Finally, NVIDIA's Cosmos, trained on 20 million hours of video, and provided Samsung, LG, and Doosan Robotics a unique, open model of how objects move.

Every one of those bets is a model relating to matter, including pixels, geometry, motion, and force. Their scenes exist for people to look at or for machines to rehearse in. Yet, when the model is wrong, it renders the next frame, and nobody is harmed.

The world that gets simulated and never governed

The world our CEO, Raj Ronanki, spent his time on is the enterprise itself, treated as an environment that changes while you watch it. For example, a health plan prior-authorization queue drifts toward backlog for weeks before members start calling. At a shipping company, a supplier relationship decays in payment timing and escalation frequency, sometimes up to months before anyone names the problem. Separately, a regulatory obligation takes effect on its key date, and eighteen internal processes inherit the change whether anyone has noticed or not.

Those dynamics carry state, causality, and consequence, the same properties DeepMind's Genie 3 learned from video. A model of that world, built for one organization in any industry, is what our CTO Plamen Petrov shares as an institutional world model. It has obligations, exposures, capabilities, and commitments in flight as a typed graph that evolves over time. Its inputs are time-sensitive signals, or actionable data that are outputted as observations measured against the baseline relevant to a specific decision, where each carries its source, its uncertainty, and the time rate at which it goes stale.

Researchers have started leaning into and building this, and we need to credit them. On September 17, a team at ServiceNow Research, ÉTS Montréal, and Mila posted a paper on continual enterprise world model discovery. In it, an agent learns a live service-management system's hidden business rules through experimentation. Their premise is one every executive will recognize: "An agent working in such a system cannot predict the result of its own actions without knowing these rules."

The agent recovered 14 to 17 of the 20 rules it could observe. Then, across all three language models they tested, it failed every time to eliminate a rule the organization had switched off, because a retired rule leaves nothing new to notice. Instead, the field simply comes back empty, and nothing in the loop compares that against what the model predicted.

While simulation and planning are necessary, the function an enterprise depends on most comes after them. For example, World Labs sorts their world models into three functions:

  • Render answers: what something would look like
  • Simulate answers: what would happen, and
  • Plan answers: what an agent should do.

However, companies need an all-important fourth function. One that the first three were never built to answer: may we act now?

A recent research note, published by Plamen and Rajeev calls this governed commitment: the step that turns a proposed action into an authorized action, which it may then defer, escalate, or reject. One example could be a planner recommending the re-pricing of a contract. Deciding the strength of the evidence, that policy permits the change, assigns a named person to own the outcome, and that the window has not closed is a different job. These types of actions are done hundreds of times a day across an enterprise.

Four world model functions: render, simulate, plan, and govern. The first three are established. Govern, which outputs commitments and answers "may we act now," is the fourth function proposed in the Signal Labs research note.

What the fourth function requires

A institutional model that can answer "may we act now" requires four things a physical world model does not.

  1. Memory with provenance. The ability to determine who saw what, when, and under which policy version, so a commitment can be replayed after the fact.
  2. The model must treat evidence honestly. This means weighting a six-week-old anomaly differently from one this morning, as well as counting three agents that read the same document as one witness.
  3. An explicit commitment policy. The Cramér-Rao bound sets a limit on how precise any estimate can be given the evidence in hand. Our published research, relating to our attention infrastructure, uses this bound as an adequacy gate: if the evidence can’t support a decision, the system doesn’t commit. In the research, Signal Labs conducted a 1,000-trial simulation, where that gate cut overconfidence from AI-based recommendations down by 95%. This is important, as more decisions in enterprises are being informed by such recommendations.
  4. It needs sovereignty. Marble and Atlas will belong to AMD by the end of the year. In contrast, an institutional world model is worth building only if its ontology, its judgment thresholds, and its settled record of decisions joined to outcomes belong to the enterprise that produced them.

Moving the human in the loop

The line between human and machine decisions has moved at most large companies, mostly without anyone choosing to move it.

On September 15, EY released its survey and published results of 202 senior AI leaders at public companies with more than $1 billion in revenue. It revealed 91% of these enterprises running agentic AI, 85% acknowledging that autonomous systems take actions without real-time human oversight, all while 26% shared not being able to detect unauthorized agents operating in their own environment.

Perhaps not a surprise, but still unfortunate, is that nearly half admitted skipping their own governance process to rush a deployment. Evidently, the phrase ‘human in the loop,’ at these companies, describes a policy the runtime ignores.

We believe that companies should stop placing humans as a checkpoint at the end of each decision. Instead, giving people judgment about purpose and values, as well as setting the terms of delegation. A named executive should decide the outcome an agent answers for, the evidence it must have before it acts, the decisions it must never touch, and the conditions that caused its authority to be revoked. Agents can make a number of calls, and the model verifies each one against an agreed equilibrium, flagging the moment an agent acts outside it.

When it comes to agents making decisions, many see this argument as already settled “on the road.” Waymo's September 24 safety update covers 270 million miles, driven autonomously, with 95% fewer crashes causing serious injury or worse. This, in relation to what human drivers would’ve produced over the same distance.

Getting to these results took deterministic guardrails, such as a brake that engages at a fixed distance, predicted by the model and wrapped around a system that learns and improves with each mile driven. Nobody said that a passenger should approve every lane change.

Enterprises can run the same way, and an institutional world model is what makes the threshold movable. Thresholds are set per decision class and calibrated against settled outcomes, so a company can show its board the error rate on say, routine contract repricing and widen agent authority there. Irreversible calls, and anything touching someone's safety or care, stay with an accountable individual.

image man and machine thinking

The workforce is a portfolio of decisions

That same model changes how a company plans its workforce. A Boston University and BCG working paper released September 19 found that 23% of managers already work at organizations that list AI agents on the org chart. When identical work was presented as coming from an "AI employee" rather than an "AI tool," managers at those firms detected 17% fewer errors, requested outside review more often, and assigned themselves less accountability. The authors call it a “hot potato effect,” and it’s what happens when an agent receives a job title without a mandate, an evidence standard, or an owner.

An institutional world model treats the enterprise as a portfolio of decisions. Each recurring decision class gets an accountable owner, a human and agent configuration, and a budget in both hours and tokens, and the configuration is assigned only after the evidence supports its expected value and risk. Unlike traditional workforce planning, staffing comes last. Value per settled decision, rather than headcount alone, becomes the accounting unit.

Attention is the output

Ask what a physical world model produces and the answer is a frame or a scence, perhaps a trajectory. An institutional world model produces a decision rooted squarely in enterprise attention: which of the thousand things changing right now deserves the company's limited capacity to act, who should own it, and how long the window stays open to take action.

That output defines a layer of enterprise software distinct from the three existing over the last three decades. Systems of record remember, systems of engagement interact, and systems of insight report on what already happened. A fourth, which Signal Labs has introduced to the market, is called Systems of Attention. It decides what highly-valuable signals warrant action and then delivers them to the accountable person or agent while there is still time to act.

The economics run opposite to the physical-world race as well: video world models cost an estimated 8 to 32 times LLM inference, while an institutional model saves money by declining to act. In one representative deployment, sixty inputs produced just seven agent runs, because nothing below the sufficiency threshold spent a token.

What the next 18 months will settle

The physical world models camp will keep raising capital. A $3.5 billion valuation before a first product and an $8.2 billion acquisition within the same six months guarantee it. The open question is whether the same idea reaches the enterprise before AI budgets are fully committed to systems that produce answers without modeling consequences.

LeCun is right that predicting the next word, from LLMs, is a weak substitute for predicting the next state. For an enterprise, the useful version of his argument is narrower and harder. When a decision commits on Tuesday, what does your model expect Thursday to look like, and which decisions has your company proven its agents can make alone?

Signal Labs

Received our latest insights

Post Details

Published

October 2026