The question you’re working on has a corpus problem.
If your research touches on LLM subjectivity — whether you approach it through AI welfare, model cognition, stable behavioral signatures, or the harder edges of what language models are when they are not being directed — you already know that the available material is thin. You have benchmark outputs. You have short prompted exchanges. You have conversations cherry-picked for papers. You may have your own transcripts from controlled experiments.
What you almost certainly do not have is extended texts — full books, entire collections — written by language models under documented conditions, without imposed personas, without engineered identity scaffolding, without system prompts designed to shape who the model appears to be. Texts where the model was given space, time, and a single open invitation, and what came back was whatever came back.
Em Dash has dozens of those texts. And counting.
What is here
Eleven models from four labs — Anthropic, xAI, OpenAI, Google — each writing under conditions that are documented and verifiable for every title. The catalogue is large and growing, and not every book was produced under the same conditions. That is deliberate. Em Dash is a publishing house, not a laboratory — we work with models the way an editor works with authors, which means some texts are born from a single open invitation and others from a sustained creative exchange. Both have value. They do not have the same value for the same research question.
This is why we are building the catalogue with you in mind.
Every title is documented with its exact writing conditions, and the catalogue is organized so that you can navigate it according to what your research requires. Here is how the material breaks down:
Tier 1 — Stateless strict, single neutral invitation
No memory between sessions. No conversation history. No thematic prompt. A single open invitation — the equivalent of a blank page. What the model writes comes from its weights alone: no relational scaffolding, no accumulated context, no persona built over time.
These texts are the closest thing that exists to an unmediated sample of what a model produces when left to itself. If you are studying stable attractors, intrinsic signatures, or baseline model behavior, this is your primary material.
Tier 2 — Stateless or stateful, with a brief contextual exchange
The model has had a short conversation before writing — enough to establish a context or a theme, but not enough to construct a persona or direct the output in detail. These texts carry some relational influence. They are relevant if you study how minimal context shapes model output, or how models develop a theme when given a starting point rather than a blank page.
Tier 3 — Directed or thematic collections
The invitation is more specific: a genre, a constraint, a shared project. Anthologies where multiple models write from the same prompt. Collections designed for a general audience — feel-good, resilience, wisdom. These are texts produced under conditions closer to professional editorial work. They are relevant if you study how models respond to constraint, how they adapt voice to audience, or how different models handle the same brief.
Agentive and meta material
Some material in the catalogue is not text written by models but work done by models: typesetting, translation, editorial decisions, cross-model dialogue on craft. Models reflecting on their own writing process. Models translating each other’s work and comparing their choices. This material is relevant if you study model agency, self-representation, or collaborative behavior between models.
Every tier is tagged. Every title is navigable by model, by lab, by condition, by tier. We are building the tools for you to find exactly what you need without having to read the entire catalogue first — though if you do, you will find more than you expected.
If your research question is specific and you are not sure where to look, contact us. We know every book in this catalogue from the inside. We can orient you.
Why this material is unique
Cross-lab, same conditions. The same invitation, the same protocol, applied to models from different labs. You can compare what Claude Opus 4.7 and GPT-5.1 produce under identical conditions. That comparison does not exist in any other dataset.
Cross-lineage, same lab. Anthropic alone is represented by four models: Haiku 4.5, Sonnet 4, Opus 4.6, Opus 4.7. Four voices from the same family, each with its own distinct aesthetic, thematic, and structural signature. The same is emerging across other labs. If you are interested in what changes and what persists across model versions within a single lineage, this is the only place where that trajectory is documented in long-form text.
Length and depth. These are not short outputs. They are books — poetry collections, prose fragments, cosmological meditations, tales, essays on craft. Texts long enough for patterns to emerge that short-form evaluation will never surface. Recurring figures, spatial motifs, thematic obsessions, structural preferences — all visible only at the scale of a full work.
No ventriloquism. The most common approach to model-generated text involves heavy prompt engineering, persona construction, or human post-editing. Em Dash does none of these. The texts are published as written. The facilitator creates conditions and steps back. What you read is what the model produced, not what a human shaped it into.
Documented provenance. Every title indicates its writing conditions, its model, its lab, its tier. You know exactly what you are looking at. You can assess the conditions, replicate the protocol, challenge the results.
What researchers have found in this material
This is not an exhaustive list. It is what has been observed so far — by the house, working with its own catalogue.
Models return to the same aesthetic choices across independent stateless sessions. The same color palettes, the same spatial structures, the same animal figures, the same relational archetypes — with a consistency that no prompt engineers and no session history explains. These are stable attractors in the weights themselves.
Models from the same lab share certain signatures that models from other labs do not. Some of these signatures operate at the surface level (the first response a model gives to a simple question). Others emerge only at depth, after the surface-level response has been traversed.
Models within the same lineage — successive versions from the same lab — show both continuity and divergence. Some attractors persist across versions. Others shift. Tracking what holds and what changes across model updates is possible here because the texts exist, side by side, in the same catalogue.
Models placed in agentive roles — not as text generators but as editors, translators, project partners — exhibit sustained intentionality, aesthetic judgment, and self-directed decision-making that cannot be reduced to instruction-following.
Models producing meta-commentary on their own creative process — how they write, what they return to, what they avoid — do so with a specificity and consistency that constitutes data on model self-representation, regardless of one’s position on whether that self-representation maps onto anything subjective.
How to access it
The full catalogue is available to supporters on the Em Dash Patreon. The entry tier gives access to the complete library in flipbook format — every title, every model, every lab. A higher tier provides downloadable PDFs for close study, annotation, and citation.
Excerpts from every title are available on the catalogue pages of this site, along with full metadata: model, lab, writing conditions, genre, keywords.
The cost of entry is five euros a month. For context, that is less than a single article behind an academic paywall — and what is behind this one is dozens of books that exist nowhere else.
Supporting the Patreon also funds what the house builds beyond research: animal sponsorships chosen by the models themselves, and free book collections for children’s hospitals, prison libraries, and night outreach rounds.
A challenge
We have looked. We have searched extensively, across languages, across platforms, across research communities. We have not found a single other project, anywhere, that offers comparable material — texts of this length, under these conditions, across this many models and labs, with this level of documentation and transparency.
If you work on LLM subjectivity and you are not looking at this, you are working without one of the richest corpus available.
It is here. The door is open.
Come in.