Ask a language model to explain a difficult subject and it can give you a clear, patient answer. Ask it to defend a mistake and it may do that just as fluently. The reader has to work out which happened.
That matters when the purpose is learning. Someone approaching a subject for the first time may be least able to spot the error, and most likely to trust a confident explanation.
01 / How text is produced
Many widely used language models are autoregressive Transformers. They break text into tokens—words, parts of words, or punctuation—and learn to predict a next token from the preceding context. During generation, that process repeats as the response grows. The GPT-3 paper describes one influential example of this approach. [1]
Within the Transformer, attention lets a token’s representation draw on other positions in the sequence. Layers of computation develop those representations before the model produces an output. The original Transformer paper introduced attention-based sequence processing; modern language models build on that family of ideas. [2]
Illustrative completion. No model probabilities are shown.
Prediction is a powerful training signal. It can produce abilities that go well beyond finishing a familiar phrase. GPT-3, for example, demonstrated translation, question answering, and generated news text that evaluators struggled to distinguish from human writing. [1]
02 / Where truth enters
Learning how language is used gives a model access to a great deal of information. But the token prediction objective does not, on its own, require an answer to carry verified evidence for every claim.
Training and evaluation matter here. Research on hallucination argues that rewarding correct guesses while penalizing uncertainty can encourage models to answer when they should acknowledge a gap. An error can arrive in the same polished prose as a fact. [3]
Retrieval can help. The original retrieval-augmented generation research combined a language model with an external document index and reported improvements in factual generation over its comparison baseline. That is evidence for a useful intervention, rather than a guarantee that every retrieved passage or generated conclusion is right. [4]
Truth is not democratic.
A claim can appear in a thousand places because those places repeat the same source. Agreement can grow without new evidence. Our position is that AI for knowledge should examine what supports a conclusion, what challenges it, and whether it applies to the question being asked.
03 / The pull of agreement
After pretraining, assistants often undergo further training to make their responses more useful. Human preferences can be part of that process. But approval is an imperfect signal for truth.
Research on sycophancy found that several assistants changed their responses toward users’ stated views, and that preference judgments sometimes favored convincing agreement over correct answers. This links one failure mode to training incentives; it does not establish that every model always agrees or that attention itself causes flattery. [5]
A useful assistant needs room to disagree. It should be able to explain a correction without turning it into a confrontation, and admit uncertainty without hiding behind vague language.
The opposite of sycophancy is not automatic contrarianism. A system that rejects the consensus just to sound independent makes the same mistake in reverse: it substitutes a posture for evidence.
04 / What writing demands
It would be easy to conclude that language models cannot write. Their actual output makes that claim hard to defend. The harder question is whether a piece of writing does the intellectual work its subject requires.
A good explanation chooses what the reader needs first. It notices when two similar terms mean different things. It handles an exception that spoils a neat summary. It revises a conclusion when the evidence changes.
Those are standards for the finished work. A fluent paragraph can meet some of them and fail others. The model’s architecture alone does not settle the quality of every piece it produces.
For Saeon, writing belongs inside the larger task of making knowledge usable. We want explanations that help a reader form a sound understanding, ask a better question, and find their way back to the evidence.
05 / Our ambition
Civilization has accumulated more knowledge than any person could read. Some of it is difficult to find. Some is difficult to interpret. Much assumes a vocabulary or background the reader has never been given.
Our goal is to reduce that distance. We’re researching AI that can help people work through language, understand the claims it carries, and use what they learn.
That means pursuing systems that can distinguish an observation from an interpretation, preserve disagreement, and explain uncertainty. It also means writing that respects a reader’s time without flattening the subject.
Spinach is one place where we explore the reading experience. The larger work continues.
Knowledge is power. People deserve a way into it.
Sources & scope
- Brown et al. — Language Models are Few-Shot Learners (2020)
- Vaswani et al. — Attention Is All You Need (2017)
- OpenAI — Why language models hallucinate (2025)
- Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020)
- Sharma et al. — Towards Understanding Sycophancy in Language Models (2023)