lanox

newchar-leveltrained from scratch

I trained a model

No big pretrained brain and nobody else's voice mixed in.Just her messages.

Lanox, an anime girl with white hair and cat ears
waking her up

01 / the difference

Same text.
Two replies.

Big models are trained to be pleasant to everyone. This one was trained on a single person, so it answers the way she does — short, dry, and entirely unbothered.

any assistant

trained on everyone

how are you

I'm doing great, thank you so much for asking! 😊 How about you? I hope your day is going wonderfully!

chars
101
emoji
1
!
2
tone
polite
L

lanox

trained on one person

how are you

chars
9
emoji
0
!
0
tone
blunt

// illustrative replies, written to show the style each one is trained toward

02 / the build


The corpus is small and the compute is free, so the training is stubborn about it: learn to be fluent, then learn to be her, then get nudged toward replies a judge can't tell apart from the real ones.

  1. 01 · step 0

    From scratch

    No pretrained weights, no borrowed vocabulary. A hand-written decoder-only transformer that learns text one character at a time — starting from noise.

    herq3l;z h.. ekkvn xo@ ,,m
  2. 02 · design bet

    One model, not two

    Style and content share the same weights. Generate-then-restyle pipelines give you a fluent stranger wearing her outfit. This doesn't do that.

    herhhe yu aa wht th.. ok ok
  3. 03 · phase 1

    Pretrain

    First it learns to be fluent at all, on synthetic casual dialogue from a teacher model. A plausible guess at how people text — explicitly not her, yet.

    herHey! I'm doing great, thanks for asking. How about you? 😊
  4. 04 · phase 2

    Fine-tune

    Then it meets the real timeline, at a much lower learning rate, so her actual habits overwrite the generic fluency instead of erasing it.

    herhey whats up
  5. 05 · phase 3

    RL with a judge

    A GRPO loop samples a group of replies, an LLM judge scores them, the policy leans toward the good ones. Only her side is trainable — the other side is always real history.

    herwhy are you asking me that
  6. 06 · the leash

    KL penalty

    A divergence penalty against the frozen Phase 2 model keeps her voice from drifting into reward-hacked mush. It's allowed to improve, not allowed to become someone else.

    heridk you tell me

The judge does not reward warmth.

// if she'd be blunt, blunt scores full marks. no bonus points for
// affection she wouldn't actually send. that's the whole trick.

the training stack

  • PyTorch

    hand-written model + loop

  • char-level

    398 symbols, emoji included

  • MiniGPT

    decoder-only transformer

  • AdamW

    cosine LR, warmup

  • GRPO

    group-relative RL

  • LLM judge

    scores realism, not warmth

  • Qwen

    fallback judge

  • KL penalty

    keeps her voice hers

03 / inside the net

The whole thing,
in numbers.

Small enough to train on a free Colab T4 — and honestly small enough to run without a GPU at all. That's the point.

her whole timeline, as raw text
Characters seen
context → her reply
Training pairs
emoji = one char
Vocab size
tiny, on purpose
~15M
Parameters

she's typing

Go on. Text her.

Fair warning — she was trained to be herself, not to be nice about it.