I have been very inspired by LLMs being extremely good 1 at solving a wide variety of tasks especially robotics, video, 3D, and 4D generation. A late realization has been that almost all of vision and graphics is actually a language modeling task 2. This has gotten me thinking about if there are any limits that arise from modelling human language with current training paradigms? Without getting into a lot of detail, here modeling means learning $P(X)$.

We need to be a bit careful when attempting to answer the question because “human language” is very loosely defined. Particularly, intelligence is the process of: collecting data in innovative ways, filtering it, the training recipe, the training code, and so on. Thus, we could very well point this process at anything and say that it is a language model.

A tempting direction to try to answer the question is to try and think about the limits of the representation we learn over: English. Now, an easy way to argue about this is that learning over a human-made language goes against the bitter lesson but this is a bit difficult because at some level of the process, we are always violating the bitter lesson. To get back to limits of the representation, it is extremely hard for me to think of any obvious limits of things that: (L1) cannot be represented in English or (L2) cannot be represented by learning the structure of the language and working with any new language at runtime.

Changing the language does make some problems easier or harder for example: many problems get extremely simple if you just use the language of Calculus 3. But, in my opinion, this is not a good way to think about any hard limits of current training paradigms. We want to gauge if there may be something in the future that may be impossible to represent by doing either L1 or L2.

Because this is getting hard to answer, I want to propose a thought experiment. The thought experiment is almost analogous to VP =? VNP and P =? NP. In the way that solving VP =? VNP would not imply anything about P =? NP but it is a simpler stepping stone. This thought experiment serves the same purpose if there are any arguments about this it may be a simpler stepping stone for the original question.

The experiment is as follows:

Assume you have a new language. This language has no notion of causality. This is extremely hard to imagine because this means there would be no order, no notion of time, no notion of cause and effect, nothing to represent ideas like “before”, “after”, “if”, “then”, etc. Now let us follow the language model process and learn your $P(X)$ for this new language. And there are some problems in doing this: 4 but for now, let us conveniently forget about some technical details (though these could impact the answer). In the process of doing the language modeling, can you pick up the notions of causality? of all the instances mentioned above?

I feel the answer might be yes (strong emphasis on feel😅) but I want to understand others’ thoughts on this.

I think some experiments like Talkie LM are very similar in spirit and should be very useful analogies. Particularly, one of the earlier examples I shared here on calculus could very well be resolved by some vintage model. Someone shared this paper with me which I haven’t read yet.

Citation

Please cite this work as:

Rishit Dagli. "A thought experiment on the limits of Language Modeling". Rishit Dagli's Blog (September 2026). https://rishit-dagli.github.io/lm-limits/

Or use the BibTex citation:

@article{ dagli2026a,
  title = { A thought experiment on the limits of Language Modeling },
  author = { Rishit Dagli },
  journal = {rishit-dagli.github.io},
  year = { 2026 },
  month = { September },
  url = "https://rishit-dagli.github.io/lm-limits/"
}

References

  1. The vision-language models currently are still using vision encoders which at least undergo a little bit of vision-only training. ↩

  2. I think current demos of seeing robotics tasks, and video, 3D generation being solved by LLMs don’t really reflect the learnings from the bitter lesson. These current demos are heavily using traditional computer graphics tools like EEF inverse kinematics, blender, simulators and more. But this is not a problem, gradually these tools will be integrated in the model itself or the model itself will be able to generate these tools on the fly for e.g. by learning the harness. ↩

  3. Of course modern LLMs also learn math very well. ↩

  4. Now if we were to do this, there are some problems like what do you do about the causal attention mask, how do you assign probability independent of the serialization order, and so on. This experiment after all is also questioning the NTP objective. ↩