ExpertsOnly a slice of the model works per word
The 105B model has 128 expert sub-networks; a router picks 8 per token, so about 10B weights do the work and serving stays cheap.
MemoryCompressed attention memory
The 105B model uses latent attention, which shrinks what it must remember about earlier text, so 128K-token inputs fit in less GPU memory.
LanguagesData and tokenizer built for India
Training leaned on the 10 most-spoken Indian languages, in native script, in Latin letters and mixed with English, the way people type.
TrainingFull pipeline at home
Pre-training, supervised fine-tuning and reinforcement learning all ran on 4,096 H100s hosted by Yotta under the IndiaAI Mission.