nkthebass commited on
Commit
ff736ce
·
verified ·
1 Parent(s): a06664f

Document what each generated dataset contains and the defect it fixed

Browse files
Files changed (1) hide show
  1. README.md +27 -0
README.md CHANGED
@@ -396,6 +396,33 @@ Then two merged adapters, 1 epoch each:
396
  The `math-balanced*` generators are the ones that mattered; each fixed a specific measured
397
  defect in its predecessor.
398
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
399
  ### Three notes for anyone reproducing this
400
 
401
  * **The predecessor's recipe did not transfer.** Rounds math2/math3 ran the exact data mix
 
396
  The `math-balanced*` generators are the ones that mattered; each fixed a specific measured
397
  defect in its predecessor.
398
 
399
+ ### What each dataset contains
400
+
401
+ All arithmetic sets are generated locally by deterministic scripts with fixed seeds, so every
402
+ label is exact by construction rather than model-generated. Names are opaque on their own, so:
403
+
404
+ | Dataset | Contents | The defect it addressed |
405
+ |---|---|---|
406
+ | `math-v3`, `math-v3fix` | column scratchpad in `ones:/tens:/hundreds:` form | the original format; `fix` corrects phantom place-value columns |
407
+ | `math-short`, `math-short2` | shorter arithmetic traces | the 320M recipe's traces were too long to learn from at this scale |
408
+ | `math-balanced` | explicit **(operation x operand-length)** joint distribution | lengths sampled independently let multiplication dominate the 2-digit band and steal addition's template |
409
+ | `math-balanced2` | + operation cue, worked partial-product sums | the model asserted partial-product sums it could not do mentally |
410
+ | `math-balanced3` | + equal-width subtraction | one extra phantom column on same-width problems |
411
+ | `math-balanced4` | + explicit `Pad 9 to 009.` statements | nothing distinguished real zero-padding from a column that should not exist |
412
+ | `math-phrased` | 30 question templates per operation | the model only recognised `"What is X plus Y?"` |
413
+ | `wp-calc` | word problems whose `<think>` block **runs** the column routine | traces asserted mental results instead of computing |
414
+ | `wp-cue` | one-step problems that **name the linguistic cue** before choosing an operation | `"drops in 836 more"` was being read as subtraction |
415
+ | `math-negative` | subtraction with negative results, explicit sign decision | every generator ordered its operands, so negatives were never trained *or tested* |
416
+ | `math-mult3` | 3-digit multipliers | the corpus contained **zero** examples with a multiplier wider than 2 digits |
417
+
418
+ Ballast, unchanged throughout: `smoltalk` and `qa-distill` (general instruction data, 9-11%
419
+ of every round) prevent the model degenerating into a calculator that cannot hold a sentence.
420
+ `reasoning-v2`, `gsm8k-cot` and `reasoning-math` are word-problem and chain-of-thought sets;
421
+ the last two are ~100% `<think>`-formatted, which is where that behaviour comes from.
422
+
423
+ The generators are deterministic - same seed, same bytes - so the exact training sets are
424
+ reconstructible from the scripts without needing the data itself.
425
+
426
  ### Three notes for anyone reproducing this
427
 
428
  * **The predecessor's recipe did not transfer.** Rounds math2/math3 ran the exact data mix