Formerly Janakpur Engineering College (JEC)Affiliated to Tribhuvan University

Mistral previews Large 4, a trillion-parameter open-weight model

Mistral released Large 4 as an API preview on 6 October and plans to publish the weights by the end of October under a custom licence.

BCT

On 6 October 2026, the French AI company Mistral released a public preview of Mistral Large 4 through its API, VentureBeat and AI News reported. The model has about one trillion parameters, but only 49 billion work on each piece of text. Mistral plans to publish the model's weights by the end of October, under its own licence.

  • 49 billionparameters active for each token
  • 160+languages in the training data
  • 3,800Nvidia Grace Blackwell GPUs used for training, per AI News
  • 27 Octplanned date for the weights, as reported by VentureBeat
  • 675 billionparameters in the previous model, Mistral Large 3

What happened

Mistral AI, based in France, announced Mistral Large 4 as a public preview on 6 October 2026. Developers can use it now through the company's paid service, called an application programming interface (API). The model accepts both images and text as input and produces text as output. Mistral says it is especially aimed at coding, cybersecurity, finance, manufacturing and understanding images.

According to AI News, Mistral trained the model from scratch on 3,800 Nvidia Grace Blackwell graphics processors in its own European data centres. VentureBeat gives the figure as about 4,000 processors and says training took about two months. The training data covers more than 160 languages, including every official language of the European Union. The public preview runs on the same European infrastructure.

Mistral plans to release the weights, the trained numbers that make up the model, by the end of October. VentureBeat reports the date as 27 October, after about three weeks of testing with cybersecurity experts, partners and government authorities. Unlike some earlier Mistral models, the weights will come under a custom Mistral licence, not the open Apache 2.0 licence. VentureBeat did not report the licence terms. Mistral says the reinforcement-learning training behind the preview is still running.

The engineering behind it

Mistral Large 4 is a mixture-of-experts model. In general, such a model contains many smaller sub-networks, called experts, and a router that chooses a few of them for each token. A token is a small piece of text, often part of a word. Because only some experts run each time, the model does less computing per token than its total size suggests. Here, 49 billion of about one trillion parameters are active for each token.

The saving is in computation, not in memory. All one trillion parameters must still be stored and loaded somewhere, which is why a model this size needs a rack of expensive graphics processors to run. For comparison, VentureBeat reports that Mistral's previous large model, Large 3, had 675 billion parameters with 41 billion active, and was trained on 3,000 Nvidia H200 processors.

After the first training on text, models are usually improved with reinforcement learning. The model tries tasks and is rewarded for good results. AI News reports that Mistral's training environments include code sandboxes, web search and outside tools, with answers checked by reward models, unit tests and other methods. At a scale of 3,000 processors, Mistral says one run produces about 33 billion tokens of training data per day.

What the benchmarks show

All benchmark figures come from Mistral and have not been checked independently. The company says the model scores 82 percent on a test that asks a model to reproduce a real security flaw in open-source software and then fix it, which it says is the highest of any model. It also reports solving 93 percent of 40 exercises from Cybench, a security competition test. Mistral says the model was useful for analysing malware and writing detection rules.

On coding, the picture is weaker. Mistral reports about 62 percent on the DeepSWE software engineering test. VentureBeat points out that the live leaderboard shows several closed models near 74 percent. In a blind human rating of coding answers run with Surge AI, Large 4 ranked second of five models, behind Claude Opus 5. VentureBeat also notes that the model is not yet listed on independent leaderboards, so Mistral's claims about its rank remain provisional.

Mistral also reports results outside coding. On AutomationBench, which covers 657 business workflows in tools such as email, spreadsheets and chat apps, it reports 59.9 percent. On a public test of attacks that try to trick a model with hidden instructions, it says the model resisted 93.3 percent. AI News notes that this is a benchmark score, not proof that the model resists almost every such attack in real use.

What it means in Nepal

Open weights matter because an organisation can download them and run the model on its own servers. Mistral says it plans to support private cloud and on-premise use, so organisations can apply their own rules. For a government agency, a hospital or a university, this means sensitive data does not have to leave its own building or country. That is different from a closed model, which can only be used through the maker's service.

Nepal's National AI Policy, approved by the cabinet in 2025, plans a National AI Centre and supports building data centres, including in cold Himalayan regions, according to The Kathmandu Post. Running a model of this size would need large and costly computing hardware. For most students, the practical route is to use such models through an API, or to study and run smaller open models on ordinary hardware.

The design ideas are useful to learn regardless. Mixture-of-experts routing, the cost of storing parameters versus the cost of computing with them, and how benchmarks are chosen and reported are all things an engineer working with AI will meet. Learning to read a company's benchmark claims with care is part of the same skill.

What to study if this interests you

Artificial Intelligence, ENCT 351, in the sixth semester of BCT, teaches neural networks, learning and how models are evaluated, and the course has a full guide on this site. Distributed and Cloud Computing, ENCT 411, in the seventh semester, explains how work is spread across many processors and servers, which is how models like this are trained and served. Network and Cyber Security, ENCT 463, in the eighth semester, connects to the security tasks Mistral says the model is built for.

Words in this story

Mixture of experts
A model design where a router sends each token to only a few of many sub-networks, saving computation.
Open weights
A release where the trained numbers of a model are published so others can download and run it.
Reinforcement learning
A training method in which a model tries tasks and is rewarded for good results, so it improves over time.
Benchmark
A fixed set of test tasks used to compare the performance of different systems.

Where this comes from

Written in our own words; no sentence is copied from these reports. Researched with AI assistance on 11 October 2026; no member of faculty has reviewed it yet. If you spot a mistake, call 01-5091616 and we will correct it and say so.

Next story: 170Nepal's Election Commission says it will move toward internet voting

Last reviewed by Imperial College of Engineering. Written 11 October 2026 from the sources above.