PrismML squeezes a 27-billion-parameter model to 5.9 GB for laptops
PrismML's Ternary Bonsai 2 27B stores weights as -1, 0 or +1 and keeps 98% of the original model's benchmark score at a ninth of the size.
The AI startup PrismML released Ternary Bonsai 2 27B on 17 September 2026, a compressed version of Alibaba's open Qwen3.8 27B language model. According to PrismML, the whole model takes 5.9 GB, more than nine times smaller than the original, and keeps 98.2 percent of its average benchmark score. The weights are free to download under the Apache 2.0 licence.
- 98.2%of the original model's average score kept, says PrismML
- 54 GBsize of the original model in 16-bit format
- 1.76 bitsper weight, by PrismML's count
- 262,000tokens of context the model can handle
- 143tokens per second on an RTX 5090, PrismML's best figure
What happened
PrismML, a company that its website says grew out of a team of researchers at Caltech, has been releasing small versions of large open models under the name Bonsai. The first Bonsai 27B came out in July 2026 and kept about 95 percent of its base model's benchmark average. The new version starts from Qwen3.8 27B, a model with about 27 billion parameters made by Alibaba. PrismML did not train a new model. It took the existing one and rewrote its weights in a much smaller form.
The original model in 16-bit format takes about 54 GB, according to the model card on Hugging Face. The ternary version of the language part takes about 5.95 GB. An optional vision file of 0.63 GB adds image input, so text-only users do not need to load it. The model keeps the original's context window of about 262,000 tokens, which is the amount of text it can consider at once. PrismML says it also runs on Apple devices through a format called MLX.
PrismML reports up to 143 tokens per second on an Nvidia RTX 5090 graphics card and about 47 tokens per second on an Apple M5 Max laptop. The model card's own detailed tables show somewhat lower decoding speeds, around 120 to 130 tokens per second on the RTX 5090, and about 32 on a lower-power Nvidia L4 card. To run the files, users need PrismML's own version of the llama.cpp program, because the standard version rejects the new formats.
The engineering behind it
A neural network stores what it has learned as billions of numbers called weights. Normally each weight is a 16-bit floating-point number. Quantisation, in general, means storing each weight with fewer bits, which saves memory and often speeds up the model. The difficulty is keeping the model accurate when so much detail is removed. Many common methods stop at four bits per weight, because quality falls quickly below that.
Ternary Bonsai goes much further. Each weight becomes one of only three values: minus one, zero or plus one. The model card explains that weights are handled in groups of 128, and each group shares one 16-bit scale number that sets its size. With the scale included, the cost is about 1.7 bits per weight. PrismML's announcement gives 1.76 bits and the model card gives 1.71 to 1.72, depending on what is counted. Either way, it is close to a tenth of the original 16 bits.
The model card also mentions a mathematical step called a Hadamard rotation, applied to blocks of weights before rounding. In general terms, a rotation spreads unusually large values across many weights, so that rounding to three levels loses less information. Ternary weights also make arithmetic simpler, since multiplying by minus one, zero or plus one is just a sign change, a skip or a copy. The card notes that unpacking the packed values costs extra work, so the smallest format is not always the fastest.
The quality loss is small but uneven. On the model card's 14 tests, the full model averaged 86.32 and the ternary model 84.78. Maths and coding stayed almost level. Knowledge questions and vision tasks lost the most: one image test, MMMU-Pro, fell from 81.73 to 75.49. DataCamp's review advised checking results carefully for important reading of text from images, and said community users reported weaker recall beyond about 200,000 tokens.
What it means in Nepal
The sources do not discuss Nepal. The general point is that a model of this size no longer needs a data centre. A file under 6 GB can, in principle, be loaded on a well-equipped personal computer, and once downloaded it runs without an internet connection. The reported speeds come from expensive graphics cards and high-end laptops, so an ordinary student laptop will be much slower, and memory remains the main limit. Still, the gap between research models and personal hardware is shrinking.
For students, compression is a practical research topic. A final-year project can take an open model, compress it in different ways, and measure exactly what is lost on tasks that matter locally, such as reading Nepali text or answering questions about a course syllabus. That kind of careful measurement, with clear tables and honest limits, is the same work PrismML published. The free Apache 2.0 licence also means students can study and modify the weights legally.
Running a model locally also changes the cost and privacy picture. A cloud model charges for every token and sends each question to someone else's server. A local model costs nothing per question once the hardware exists, and the data stays on the machine. For a project that handles private records, such as student marks or medical notes, that difference can matter more than a few points on a benchmark. Weighing those trade-offs is a normal engineering decision.
What to study if this interests you
Digital Logic, ENEX 152, in the second semester of BCT and BEI, starts with number systems and binary representation, the basis for understanding how few bits a weight can use. It has a full guide on this site. Numerical Methods, ENSH 252, in the fourth semester, deals with rounding and truncation error, which is exactly what quantisation trades against size. Computer Organization and Architecture, ENCT 303, in the fifth semester of BCT, explains why memory size and bandwidth limit speed.
Artificial Intelligence, ENCT 351, in the sixth semester of BCT, introduces neural networks and how they learn, which explains what the weights in this story are and why removing detail from them affects some answers more than others. The course has a full guide on this site. A minor project in the sixth semester is a natural place to try running and measuring a compressed model.
Words in this story
- Weight
- One of the learned numbers inside a neural network that decides how strongly one signal affects the next.
- Quantisation
- Storing numbers with fewer bits than before, which saves memory at the cost of some precision.
- Ternary
- Using three possible values, here minus one, zero and plus one, instead of the two values of binary.
- Context window
- The amount of text, counted in tokens, that a language model can read and consider at one time.
Where this comes from
- PrismML, 17 Sep 2026
- Hugging Face model card (prism-ml), 17 Sep 2026
- DataCamp, 18 Sep 2026
The news itself rests on one source; any other link is background or from the same publisher. Written in our own words; no sentence is copied from these reports. Researched with AI assistance on 11 October 2026; no member of faculty has reviewed it yet. If you spot a mistake, call 01-5091616 and we will correct it and say so.




