Mistral AI CEO Arthur Mensch sat down with investor Elad Gil at Figma HQ to discuss how his team built and shipped the Mistral 7B model in four months, starting from zero GPUs and scaling to 500 GPUs. Mensch, a former DeepMind staff research scientist, co-founded Mistral with colleagues from DeepMind and Meta's LLaMA project after GPT signaled the window was open. The company launched 7B, then Mixtral 8x7B in December, and added commercial models by February, all while keeping the codebase open source and layering a closed commercial platform on top.
The conversation gets specific on org structure and product strategy. Mensch credits the speed to small, focused squads of four to five people covering data, pretraining, and infrastructure, no large committees, minimal holidays. On go-to-market, Mistral is targeting two lanes: financial services enterprises through cloud partnerships including a recently announced Microsoft Azure deal, and developer-native AI companies through their direct API platform. The Azure relationship matters because enterprises cannot easily route data through third-party SaaS, making hyperscaler distribution a real unlock.
The full transcript is worth reading for Mensch's unfiltered views on open versus closed source AI, the EU regulatory environment, where small models are headed as context windows expand, and what enterprises actually need versus what the market is selling them. The roadmap discussion, covering vertical-specific models and a nascent chat assistant called Le Chat, gives a concrete preview of how Mistral plans to move from infrastructure provider to a full enterprise AI stack.
[READ ORIGINAL →]