Project

The Theory of Large Language Models

A Bayesian account of how large language models learn: in-context learning as inference on a latent manifold, and the geometry of attention.

Related Publications