Research agenda

What we work on, and in what order

Century · Research · 2026

1.The structure of trained models

Every capability a model has is physically present in its weights, put there by an optimization process no one directs in detail. We treat trained models as objects of study in their own right — more like organisms to dissect than programs to read.

The near-term work is concrete: running open-weight models locally, probing their behavior under controlled variation, and building evaluations that distinguish knowing from pattern-matching. The questions we want these experiments to answer include: how is a single capability distributed across a network; what changes inside a model when it is finetuned; and which behaviors are properties of the architecture versus properties of the data.

2.Training and finetuning open weights

Open-weight models make a small independent lab viable: the frontier of understanding does not require the frontier of scale. Our experimental ladder runs from inference to intervention —

  • Now: local inference and evaluation across open-weight model families, establishing baselines we trust because we ran them ourselves;
  • Next: finetuning runs designed as experiments — each one built to confirm or kill a specific hypothesis about what training data does to model internals;
  • Then: training our own models, small enough to study exhaustively and large enough to exhibit the phenomena worth studying.

The discipline throughout is that training is our laboratory instrument, not our product. A run that produces a worse model but a sharper explanation is a success.

3.Reasoning from first principles

This is the problem the other two exist to serve. Current models are extraordinary at producing text shaped like reasoning; whether the underlying computation is reasoning — derivation from principles rather than interpolation over precedent — is exactly the kind of question that today cannot even be posed precisely. Making it precise is part of the work.

The long-term questions, stated plainly:

  • What would it mean, mechanistically, for a model to reason from first principles — and how would we measure it in a way that imitation cannot fake?
  • Can training objectives be designed that reward derivation over recall, so that a model’s confidence tracks the strength of its argument rather than the frequency of its answer?
  • Do such models generalize to domains with no precedent to imitate — novel structures, novel propulsion regimes, novel physics problems?

The last question is the bridge to our applied ambitions. A model that only retrieves can never design what has never existed. A model that derives, in principle, can — and megastructures and launch systems are precisely the domains where nothing worth building has a precedent.


A note on publishing

We intend to write up what we find, including negative results. Nothing is published here yet — the lab is new and we would rather publish something true than something soon.

Correspondence

Questions about the agenda: hello@century.sh