Research
LLMs are not enough for sports
Sport is not only a language problem. It is a world problem.
Large language models (LLMs) are extraordinarily useful. They can research, summarize notes, retrieve information, write plans, reason through a question, and give anyone a natural interface to an AI system.
For many industries, LLMs are proving to be revolutionary, unlocking capabilities that were previously out of reach. Where the work is fundamentally language — research, drafting, summarizing, explaining, turning a request into a plan — they compress hours of effort into seconds, lower the barrier to expertise, and give anyone a natural interface to software that once demanded training.
Sport is one of the places where that reach runs out. An LLM can help a coach reason about the game, but it cannot play the game forward. Not because the models are weak, but because the questions that decide games are not language questions.
That gap is what forces the harder questions into view: What happens next if I do this? And: What should I do because of it?
A language model can describe a plausible answer. Its training objective does not require that answer to behave like a grounded simulation of a particular sport, one whose predictions can be rolled forward, acted on, and compared with what actually happened. That is the problem we are working on right now.
Sports have reached the ceiling of language models
Across sport, language models have been adopted quickly and for good reason: they turn scattered information into answers, plans, and summaries in plain language. But it is important that we are clear about what has actually improved, because almost all of it is downstream of language. The model helps a person work with what has been said, written, and documented about the sport. In short, LLMs are always downstream of complex, human decision making: a person makes the call, and the model works with what that call already produced.
What has not changed is the part that decides the outcomes. A language model can describe a coverage in football, summarize a training block for a marathoner, or produce a lifting plan that reads well, but it cannot hold the state of a play, run it forward under a decision, and be held accountable to what actually happened. As the questions get more difficult, and closer to real-time, the effectiveness of LLMs declines rapidly. More language, better phrasing, and larger models do not add up to a grounded simulation of the sport.
That is the capacity we have reached. The next layer of value in sports AI is not more fluent text; it is a model that understands the dynamics of a specific sport well enough to imagine consequences and be measured against reality. That is what we are building, and we are excited to share our research so far.
Bringing world models to sports
Yann LeCun, one of the field's longest-standing critics of language models as a path to human-level AI, laid out the alternative to LLMs in his 2022 paper, A Path Towards Autonomous Machine Intelligence. It proposes the Joint Embedding Predictive Architecture (JEPA): systems that learn abstract representations of a world and predict how those representations change rather than reproducing every pixel. Our work is influenced by that research on predictive world models.

The basic idea is straightforward. An encoder turns an observation into a compact internal state. A transition model learns what different actions do to that state. The model can then apply those learned transitions repeatedly to imagine possible futures before taking an action.
In simplified form: observe → represent → imagine → evaluate → act → verify
The important word here is verify. In the small exact worlds we are using for research today, we know what should actually happen. That means an imagined future can be compared with reality rather than merely sounding plausible. A world model tries to learn something different: how a particular world changes. Give it a state and an action, and it predicts the state that follows.
A coach does not only want to know what a blitz is. Eventually, we want a system that can represent this situation, consider several actions, imagine how each could unfold, compare those imagined futures, and then be held accountable when reality arrives. That requires more than fluent language; it requires a model with dynamics.
What we are building toward
Our goal is a sport-specific intelligence layer that can maintain an internal state of a game or training environment, learn reusable dynamics, imagine the consequences of actions, compare possible futures, notice when reality violates its expectations, and adapt without discarding useful knowledge unnecessarily.
Language models still belong in that system. They are how a coach, athlete, or trainer can talk to it. They can retrieve information, explain its reasoning, summarize evidence, and turn a complicated model into something useful to a person.
But underneath that conversation, we believe sport needs another kind of intelligence too: a model that is accountable to the world it is talking about. That is what we are researching.
It is early. Some pieces work. Some approaches have failed. Several of the hardest pieces remain open. That is precisely why we are building the small worlds first. Before we ask a model to understand a fourth-quarter defense, a pitcher, a tennis rally, or an athlete's training state, we want to know something much simpler: When it imagines what happens next, can we trust that imagination, and does it know when it should not? That is the standard we are trying to build toward.
Below is a small example of a maze world model performing that behavior: it reads the state of the board, imagines how its actions could unfold, and replans in real time when the world disagrees with what it imagined.
The model reads the board, imagines several short futures, and commits to a single move.
Planning is not the same as prediction
One of the more useful lessons from our experiments has been that predicting the world well does not automatically tell an agent how to make a good decision.
Our current maze planner can roll actions several steps forward through a learned latent world model. A seemingly obvious strategy is to score every imagined future and choose the single numerically best one. That did not work as well as we expected.
Instead, our stronger results came from treating a region of futures as good enough. Rather than pretending that tiny differences between two model predictions are necessarily meaningful, the controller identifies a basin of promising futures and chooses from within it. That distinction matters: a model may reliably understand that several futures are all heading in the right direction without reliably knowing that one is 0.00002 “better” than another. For planning, useful structure can matter more than false precision.
When the world surprises the model
Another question we have been studying is what an agent should do when reality disagrees with its prediction. The obvious machine-learning answer is: prediction was wrong → retrain the model.
Our experiments suggest that this is sometimes exactly the wrong response. In one controlled family of maze experiments, the underlying transition behaviors remained familiar but the relationship between external commands and those known behaviors changed. Retraining the world model improved some prediction measurements, but it did not reliably restore good control and could interfere with knowledge that already worked.
A much simpler approach worked better: first ask whether the surprise can be explained by a different configuration of things the model already knows. When it could, the model was able to re-identify that structure without rewriting its learned dynamics.
We have since tested that idea across the complete family of command mappings in that experimental world, and we have shown that a lightweight belief over known configurations can also recover when the active mapping changes unexpectedly during an episode.
That does not mean every new formation, athlete, rule, or strategy in a real sport is merely a remapping problem. Quite the opposite. The research question we care about is learning to distinguish: “I already understand this, but my current interpretation is wrong.” from: “None of what I know explains this. I may actually need to learn something new.” That boundary is one of the things we are working toward.
Memory is part of the problem too
A useful world model cannot live entirely in one isolated decision. It needs some notion of context: what kind of world it currently believes it is operating in, and how strongly it should carry that belief forward.
Our recent experiments have started testing this directly. Carrying a structural belief across tasks helps when the underlying regime remains stable, and the system can still recover when that regime changes.
But simply carrying the same belief forever is not enough. In our current experiments, it has not yet demonstrated the kind of recurrence behavior we would want from a true persistent configurator.
That is useful evidence. The question is no longer merely whether an intelligent system should remember. It is: what should it remember, how strongly should it remember it, and what evidence should cause it to change its mind?
Why we are starting with small worlds
Our conviction is straightforward: world models will solve AI in sports. The questions that decide games are questions about how a world changes under action, and a system that learns exactly that is not an incremental improvement over language models — it is the right shape for the problem itself.
But conviction is not evidence, so we are starting small to prove it. We are beginning with exact domains such as mazes and strategy games because their rules and transitions can be known. If a model predicts what happens after an action, we can check it. If a planner claims one strategy works better than another, we can run it. If a cheaper planning method looks promising and fails on fresh confirmation data, we can record the failure instead of quietly changing the benchmark.
We have already had several approaches that looked good initially and did not survive that process. We keep those results. None of these claims require a football field yet — in fact, starting there would make the science worse. The point is to prove the idea somewhere measurable first, then carry it onto the field.