Multilingual pre-training mixture and scaling optimization service
Help AI developers choose multilingual training-data mixtures and model sizes by empirically measuring transfer, synergy, and interference between language or other data sources.
From Y Combinator — Going In Deep On Data | YC Paper Club at 04:05
Problem: Low-resource languages have far fewer training tokens than English, while monolingual training can repeatedly reuse scarce data and overfit. Adding other languages can help or hurt, and the relationship is not reliably determined by language family or intuition alone.
For: AI developers building models for underrepresented languages, cultures, societies, or domains with uneven data availability.
Examples
- Thai had about 0.6% as many tokens as English in Madlad 400, illustrating the data constraint for a lower-resource language.
Behind this: 14 build steps · 2 tools and how each is used · how to validate demand · 3 more real examples · 4 things the video never answers.