any assistant
trained on everyone
I'm doing great, thank you so much for asking! 😊 How about you? I hope your day is going wonderfully!
- chars
- 101
- emoji
- 1
- !
- 2
- tone
- polite
newchar-leveltrained from scratch
No big pretrained brain and nobody else's voice mixed in.Just her messages.

01 / the difference
Big models are trained to be pleasant to everyone. This one was trained on a single person, so it answers the way she does — short, dry, and entirely unbothered.
any assistant
trained on everyone
I'm doing great, thank you so much for asking! 😊 How about you? I hope your day is going wonderfully!
lanox
trained on one person
// illustrative replies, written to show the style each one is trained toward
02 / the build
The corpus is small and the compute is free, so the training is stubborn about it: learn to be fluent, then learn to be her, then get nudged toward replies a judge can't tell apart from the real ones.
her (allegedly)
checkpoint · step 0
…still learning letters
01 · step 0
No pretrained weights, no borrowed vocabulary. A hand-written decoder-only transformer that learns text one character at a time — starting from noise.
02 · design bet
Style and content share the same weights. Generate-then-restyle pipelines give you a fluent stranger wearing her outfit. This doesn't do that.
03 · phase 1
First it learns to be fluent at all, on synthetic casual dialogue from a teacher model. A plausible guess at how people text — explicitly not her, yet.
04 · phase 2
Then it meets the real timeline, at a much lower learning rate, so her actual habits overwrite the generic fluency instead of erasing it.
05 · phase 3
A GRPO loop samples a group of replies, an LLM judge scores them, the policy leans toward the good ones. Only her side is trainable — the other side is always real history.
06 · the leash
A divergence penalty against the frozen Phase 2 model keeps her voice from drifting into reward-hacked mush. It's allowed to improve, not allowed to become someone else.
The judge does not reward warmth.
// if she'd be blunt, blunt scores full marks. no bonus points for
// affection she wouldn't actually send. that's the whole trick.
the training stack
hand-written model + loop
398 symbols, emoji included
decoder-only transformer
cosine LR, warmup
group-relative RL
scores realism, not warmth
fallback judge
keeps her voice hers
03 / inside the net
Small enough to train on a free Colab T4 — and honestly small enough to run without a GPU at all. That's the point.
she's typing
Fair warning — she was trained to be herself, not to be nice about it.