DataBench
Rather than training one model, we trained multiple models, each on a different dataset. Everything else was kept the same, so that the only changing variable would be the dataset itself.
Model Architecture
- Base Architecture:
LlamaForCasualLM - Tokenizer:
Harley-ml/Dillionv2-1.3M - Transformers Version:
5.13.1 - Hidden Size:
128 - Vocab Size: 2564
- Number of Layers:
8 - Number of Heads:
4 - Number of KV Heads:
2 - Intermediate Size:
344 - Head Dim:
32 - Max Position Embeddings:
256 - RoPe Theta:
2500.0 - Tie Word Embeddings:
true - Hidden Activation:
silu - MLP Bias:
false - Initializer Range:
0.2 - RMS Norm Eps:
1e-06 - Pretraining Tp:
1 - Use Cache:
false - Total Parameters:
1,780,352
Training Setup
- Epochs:
1 - Max Steps:
-1.0 - Batch Size:
400 - Sequence Length:
256 - Gradient Accumulation:
2 - Gradient Clipping:
1.0 - Gradient Checkpointing:
true - Learning Rate:
2.5e-3 - Eval Split:
0.00165 - Weight Decay:
0.01 - Optimizer:
AdamW - AdamW Betas:
(0.9, 0.95) - AdamW Eps:
1e-8 - Scheduler:
WSD - WSD Warmup Ratio:
0.015 - WSD Stable Ratio:
0.78 - WSD Decay Ratio:
0.20 - WSD Minium LR Ratio:
0.0 - WSD Number of Cycles:
0.5 - DType:
float16 - Torch.Compile:
true - DataLoader Workers:
2 - Seed:
311
Results
Accuracy is normalized by length and shown as a percentage.
| Dataset | ARC-Easy | HellaSwag | PIQA | Avg ↑ |
|---|---|---|---|---|
| FineWeb | 27.82% | 26.89% | 52.83% | 35.85% |
| DCLM-1.0-Baseline | 29.08% | 27.06% | 52.23% | 36.12% |
| DOAB | 31.14% | 28.00% | 52.18% | 37.11% |
| Project Gutenberg | 25.84% | 24.63% | 50.21% | 33.56% |
| EOT-2004-Raw | 29.08% | 27.54% | 51.31% | 35.98% |
| Wikipedia | 28.28% | 27.61% | 51.14% | 35.68% |
(Will add ArithMark-3.0 later)
Notice
This is a work in progress and is currently not completed. By the end of this project, we aim to have tested over 50 datasets.
License
Apache 2.0.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support