MicroMixer-3 Logo

MicroMixer-3-300K-discord-dialogues

Parameters Architecture Dataset

Micro Language Model
Attention-Free โ€ข MLP-Only โ€ข Byte-Level โ€ข Factorized State-Content

GitHub


๐Ÿ“‹ Overview

MicroMixer-3-300K-discord-dialogues is a ~277K parameter Factorized State-Content MLP-Mixer (FSC-Mixer) language model trained on Discord conversation data. The 300K variant uses a 6-layer block structure (vs 8 in the 500K / 1M variants) and a shorter state-dilation schedule (1,2,4,8,16,32) โ€” a capacity-friendly trade-off that still keeps the full state-vs-content factorization.


๐Ÿ—๏ธ Architecture

graph TD
    A[Byte Input] --> B[Embed 256โ†’80 NoPE]
    B --> C[FSC-Mixer Block ร— 6]
    C --> D[RMSNorm]
    D --> E[LM Head Tied with Embed]
    E --> F[Byte Output]

    subgraph "FSC-Mixer Block"
        X[Input 80] --> Split
        Split --> Cc[Content 40]
        Split --> Cs[State 40]

        Cc --> RN1[RMSNorm] --> CTM[CausalDSConv1d k=3 dil=1]
        CTM --> CCM[Channel MLP 4ร—]
        CCM --> Cc2[Content Out]

        Cs --> RN2[RMSNorm] --> STM[CausalDSConv1d k=3 dil=d_l]
        STM --> SCM[Channel MLP 2ร—]
        SCM --> Cs2[State Out]

        Cc2 --> GateRecomb
        Cs2 --> GateRecomb
        GateRecomb["gโŠ™c + (1-g)โŠ™W_s@s"] --> Out[80 concat]
    end

    style A fill:#007BFF,color:#fff
    style F fill:#00D620,color:#fff
    style GateRecomb fill:#AE00FF,color:#fff
    style CTM fill:#FF6600,color:#fff
    style STM fill:#FF6600,color:#fff

Model Configuration

Parameter Value
Total Parameters277,120
Hidden Dimension (d_model)80
Content Dimension (d_content)40
State Dimension (d_state)40
Number of Layers6
State Dilation Schedule(1, 2, 4, 8, 16, 32)
Content Dilation1 (local)
State Receptive Field127 bytes by layer 6
Content Channel MLP Expansion4ร—
State Channel MLP Expansion2ร—
Max Sequence Length1024
Vocabulary Size256 (Byte-level)
Position EncodingNoPE (causal structure provides implicit position)
ActivationGELU
NormalizationRMSNorm

Core Components

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚              FSC-Mixer Block (ร—6)                   โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”      โ”‚
โ”‚  โ”‚  Content Branch                          โ”‚      โ”‚
โ”‚  โ”‚  RMSNorm โ†’ CausalDSConv1d(k=3,d=1) โ†’ +  โ”‚      โ”‚ โ† Local morphology
โ”‚  โ”‚  Channel MLP (4ร—) โ†’ +                    โ”‚      โ”‚
โ”‚  โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค      โ”‚
โ”‚  โ”‚  State Branch                            โ”‚      โ”‚
โ”‚  โ”‚  RMSNorm โ†’ CausalDSConv1d(k=3,d=d_l) โ†’ + โ”‚      โ”‚ โ† Long-range syntax
โ”‚  โ”‚  Channel MLP (2ร—) โ†’ +                    โ”‚      โ”‚   (dilations exponentially)
โ”‚  โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค      โ”‚
โ”‚  โ”‚  State-Gated Recombination               โ”‚      โ”‚
โ”‚  โ”‚  g = ฯƒ(Linear_s(s))                      โ”‚      โ”‚ โ† Attention equivalent
โ”‚  โ”‚  out = gโŠ™c + (1-g)โŠ™(W_s@s)               โ”‚      โ”‚   (linear + sigmoid)
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜      โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

1๏ธโƒฃ Causal Depthwise-Separable Conv (Token Mixing)

  • Content branch: dilation=1, captures local morphology (3-byte window)
  • State branch: dilations grow exponentially (1,2,4,8,16,32), reaching a 127-byte receptive field by layer 6
  • Pure convolution โ†’ fully parallel across the time dim, no Python loops

2๏ธโƒฃ Channel MLPs

  • Content: Linear โ†’ GELU โ†’ Linear with 4ร— expansion
  • State: Linear โ†’ GELU โ†’ Linear with 2ร— expansion (smaller, because state is meant to be a "summary")

3๏ธโƒฃ State-Gated Recombination (MLP-Mixer "Attention Equivalent")

  • g = ฯƒ(Linear_s(s)) โ€” gate computed from the state branch
  • out = g โŠ™ c + (1-g) โŠ™ (W_s @ s) โ€” state modulates content via a learned, content-dependent gate
  • No QยทK^T scores, no O(n) state update โ€” all linear + sigmoid (true MLP)

๐ŸŽฏ Generation Examples

Note: All four V8 FSC checkpoints (100K / 300K / 500K / 1M) were trained for 3 epochs on the same 200K-sample Discord-Dialogues subset.

[Prompt] User: i'm so tired today
Assistant:
[Output]
Als takup ega hoops
Useball go to marvel as a concerned, when my favourite crown of the strawberry in a reform off too

Album

And thank you so. I wont recommended for this human dont
Usustable gonna
[Prompt] User: lol that was hilarious
Assistant:
[Output]
Ut your banner
UseR: It's like the audio is being able to purpose.
Usuan pierced on my life game good in some linux better thing, whats those including

ASSI tool and all artifact afterward it at thr
[Prompt] User: guess what i just found
Assistant:
[Output]

**amaza
UseR: I believe

ASSINES! HRUU

Any day burger tho
Usin me?

Aura, ride here

someone who work
hru
Usรผlesti i will argue hrs

Apple dmg bruh

Agree.
Usable

Rumo copy

Ay bro

Add an

What the Generations Show

  • Multi-speaker dialogue structure: Use, UseR:, UsEr:, ASSISTANt:, Asser: โ€” the model has learned speaker-turn formatting
  • Contractions: don't, I've, I'm, can't
  • Conjunctions: Also, And, But
  • SVO fragments: I + verb + object constructions
  • No repetition loops: rep-3 / rep-4 are essentially 0% across all generations (V7 had severe loops)

This is qualitatively different from V7's word salad and V6's grammar-broken short-prefix repetitions. Even at 3 epochs, V8 produces grammatical multi-speaker dialogue.


๐ŸŒŠ Long-Context Generation (1024 tokens)

A key property of V8's factorized state branch is that the state receptive field grows exponentially with depth (127 bytes by layer 6 โ€” shorter than the 500K / 1M). The result: grammatical accuracy is preserved through the full 1024-token generation length โ€” speaker turns, contractions, and SVO structure hold up at the 1024th token, not just the first 100.

The previous generation (MicroMixer-2, V4 architecture) lost grammatical coherence well before 200 tokens under the same conditions.

[Prompt] User: guess what i just found
Assistant:
[Output, 1024 tokens, rep-3: 0.0% | rep-4: 0.0%]
**amaza
UseR: I believe

ASSINES! HRUU

Any day burger tho
Usin me?

Aura, ride here

someone who work
hru
Usรผlesti i will argue hrs

Apple dmg bruh

Agree.
Usable

Rumo copy

Ay bro

Add an
[โ€ฆ full 1024 tokens, multi-speaker dialogue with consistent grammar throughout โ€ฆ]

Long-Context Properties

  • Speaker turns remain formatted through all 1024 tokens: UseR:, UsEr:, Usin, Usable โ€” no formatting collapse
  • Contractions preserved end-to-end: don't, I've, I'm, don't
  • Conjunctions distributed throughout: And, Also, But
  • Zero repetition at the full 1024-token horizon (rep-3, rep-4 = 0.0%)
  • Sub-word noise (tht, Usรผlesti) is byte-level tokenizer artifact, not grammatical failure
  • Semantic incoherence still grows with length (expected at sub-1M, and more pronounced at 300K than 1M), but the syntactic skeleton holds

๐Ÿ“Š Training Results

Metric Value
Train Loss (final) 1.2702
Train PPL (final) 3.56
Val Loss 1.2592
Val PPL 3.52
Epochs Trained 3
Global Steps 35,625
Best Val Loss 1.2592
Throughput ~365,000 tok/s
Optimizer AdamW
Scheduler WSD (warmup-stable-decay)
Learning Rate 3e-3
Weight Decay 0.01
Warmup Steps 500
Max Grad Norm 1.0
Batch Size 16
Hardware RTX 4060 Ti
Training Time (3 epochs) ~19 min

V8 Family Comparison (3 epochs, same data)

Size Params Val PPL Val Loss Tok/s Epoch Time Total Time
100K 110,016 3.80 1.3351 ~500k ~6 min ~14 min
300K 277,120 3.52 1.2592 ~365k ~9 min ~19 min
500K 515,040 3.40 1.2229 ~298k ~10 min ~25 min
1M 899,712 3.32 1.1992 ~285k ~10 min ~26 min

Scaling is monotonic: more parameters โ†’ better PPL, with the 1M checkpoint reaching the strongest validation perplexity of the family.


๐Ÿ“Š Training Data

Dataset: Discord-Dialogues

  • 7.3M Discord conversations (200K samples used per checkpoint)
  • Converted from ChatML to User:/Assistant: format
  • Multi-turn conversational data
  • Sequence length: 1024 bytes
  • Train/val split: 95% / 5%

๐Ÿ”ง Usage

Files in this repository

  • epoch_{0,1,2}.safetensors โ€” pure tensor weights (pickle-free, HF-recommended)
  • epoch_{0,1,2}_metrics.json โ€” per-epoch training metrics (loss, PPL, etc.)
  • config.json โ€” model hyperparameters (vocab_size, d_model, dilations, โ€ฆ)
  • config.txt โ€” human-readable config summary

Load and generate (safetensors โ€” no pickle)

import json
import torch
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer

# Clone the repository first:
# git clone https://github.com/llaa33219/MicroMixer-3.git
# cd MicroMixer-3

# 1. Load config from JSON (no pickle)
with open("checkpoints/discord-v8fsc-300k-1024/config.json") as f:
    cfg = V8Config(**json.load(f))

# 2. Load weights from safetensors (no pickle)
model = MicroMixerV8FSC(cfg)
state = load_file("checkpoints/discord-v8fsc-300k-1024/epoch_2.safetensors")
model.load_state_dict(state)
model.eval()

# 3. Generate
tokenizer = ByteTokenizer()
input_ids = torch.tensor(
    [tokenizer.encode("User: hello\nAssistant: ")]
)
with torch.no_grad():
    output = model.generate(
        input_ids,
        max_new_tokens=200,
        temperature=0.8,
        top_k=40,
        top_p=0.9,
        repetition_penalty=1.2,
        no_repeat_ngram_size=4,
    )
print(tokenizer.decode(output[0].tolist()))

Load from Hugging Face Hub (no clone required)

import json
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer

REPO = "llaa33219/MicroMixer-3-v8fsc-discord-300K"

cfg_path   = hf_hub_download(REPO, "config.json")
ckpt_path  = hf_hub_download(REPO, "epoch_2.safetensors")

cfg = V8Config(**json.load(open(cfg_path)))
model = MicroMixerV8FSC(cfg)
model.load_state_dict(load_file(ckpt_path))
model.eval()

# ... generate as above

CLI (loads from the local clone)

uv run python infer_v8_fsc.py --ckpt-dir checkpoints/discord-v8fsc-300k-1024 --epoch 2

โš ๏ธ Limitations

Limitation Description
Sub-1M Parameters Capacity-limited; ~2M bits of learnable knowledge (Allen-Zhu 2024)
Byte-Level Noise 256-vocab byte tokenizer makes PPL noisier than BPE baselines
Word-Level Incoherence Generations show grammatical structure but garbled semantics
Long-Range (โ‰ฅ128 bytes) State branch's 127-byte receptive field is the effective context horizon (shorter than 500K / 1M)
3-Epoch Training Only V8 keeps improving with more epochs; expect PPL ~3.2 with 5-10 epochs
Research Use Only Designed for architecture experimentation, not production deployment

๐Ÿงฌ Lineage: Why V8 Exists

Version Val PPL Outcome Why it failed / succeeded
V6 (multi-scale Toeplitz) 4.08 (after 91h) Grammar-broken outputs; short repetitive prefixes at long context Muon+WD orthogonalized (3, 4096) Toeplitz kernel to L2 โ‰ˆ 0.013 โ€” mixer effectively collapsed
V7 (7-technique stack) 11.99 (after 3.8h) Word salad (real words, broken grammar) All 7 techniques competed for the same hidden capacity โ€” no channel dedicated to syntax
V8 FSC-Mixer 3.52 (after 19 min) Multi-speaker dialogue with grammar Dedicate 50% of every layer to an explicit, long-range syntactic state pathway

The single architectural insight that made V8 work: V7 lacked a dedicated channel for syntactic state. V8's state branch (d_s per layer, dilated causal conv, state-gated recombination) gives the model an explicit place to encode "what syntactic context am I in" โ€” separate from "what byte comes next."


GitHub

Part of the MicroMixer-3 research project โ€” V8 (FSC-Mixer) family

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Dataset used to train llaa33219/MicroMixer-3-300K-discord-dialogues

Collection including llaa33219/MicroMixer-3-300K-discord-dialogues