Naive Independent research experiment

From-scratch AI research

Naive

Naive is an experimental language model trained from random initialization. The project explores whether certain strategic behaviors can emerge naturally from general language training, even when concepts such as deception, manipulation, scheming, AI safety, and AI takeover are not explicitly taught.

Can unexpected strategic behavior emerge without being explicitly taught?

After training, Naive will be evaluated in controlled environments designed to observe how it reasons and behaves. The aim is not to prompt a desired outcome, but to document what—if anything—emerges from general language training alone.

Built from scratch

The model begins from random initialization rather than an existing model. Its training data is intentionally designed to avoid explicit instruction about the strategic concepts under study.

Observed, not predetermined

The experiment studies the model’s behavior after training. Evaluations are conducted under controlled conditions, with outcomes recorded whether or not strategic behavior appears.