Formerly Janakpur Engineering College (JEC)Affiliated to Tribhuvan University

Anthropic releases Claude Opus 5.5 at lower prices

Claude Opus 5.5 launched on 22 September at $4 and $20 per million tokens, with Anthropic claiming top-tier coding results at about 40% lower running cost.

BCT

Anthropic released Claude Opus 5.5, the first model in its 5.5 family, on 22 September 2026, according to the company and TechCrunch. Anthropic says it performs at about the level of its larger Fable 5.1 model on most work while costing about 40 percent less to run than Opus 5. API prices fell 20 percent, to 4 US dollars per million input tokens.

  • $4 / $20API price per million input and output tokens
  • 66.4%Terminal-Bench 4.0 score, against 52.3% for Opus 5
  • 680,000lines in a code migration a tester finished in under a day
  • 30%+faster output than Opus 5, Anthropic says

What happened

Anthropic sells its Claude models in three tiers. Opus is the most capable and most expensive, Sonnet sits in the middle, and Haiku is the fastest and cheapest. TechCrunch noted that Opus 5.5 arrived two months after Opus 5, which launched on 24 July. Anthropic said Sonnet 5.5 and Haiku 5.5 would follow in the coming weeks. The new model is available through Anthropic's own apps and API and through the Amazon, Google and Microsoft cloud platforms.

The new prices are 4 dollars per million input tokens and 20 dollars per million output tokens, down from 5 and 25 dollars for Opus 5. Reading from the cache, where a repeated part of a prompt is stored and reused, fell from 0.50 to 0.20 dollars per million tokens. Anthropic says the model also produces output more than 30 percent faster than Opus 5. TechCrunch linked the speed to the lower amount of computing needed to serve the model.

Anthropic's results show large gains on coding tests. On Terminal-Bench 4.0, which measures multi-step work in a computer terminal, Opus 5.5 scored 66.4 percent against 52.3 percent for Opus 5 and 55.8 percent for Fable 5.1. Anthropic reported that one early tester completed a 680,000-line code migration in under a day, and that an audit and fix of a 200,000-line codebase took under three hours, against more than 20 hours for Opus 5. TechCrunch added that Anthropic says the model now uses less jargon and puts the most important information at the start of its replies.

The engineering behind it

Most of the reported gains are about efficiency rather than raw ability. Like several recent models, Opus 5.5 lets developers choose an effort level, from low to maximum, which controls how much the model reasons before answering. More effort usually means better answers but more tokens and higher cost. Anthropic's main claim is that Opus 5.5 at its default medium effort matches or beats older models running at maximum effort, at a fraction of the cost per task.

Several customers quoted by Anthropic described the same pattern in their own tests: fewer steps, fewer tool calls and fewer tokens to finish the same job. GitHub, for example, said the model solved more terminal tasks than Opus 5 in under half the steps. Stripe said one session directed a dozen others to rebase 40 linked code changes, and all 40 passed its automated tests. These are company statements, not independent measurements.

Anthropic itself added cautions about its numbers. It published a margin of error of about 2.6 points on the Terminal-Bench result, and said that at this level of capability, benchmark margins are a less reliable guide than before. It also said that the gap to Fable 5.1 looked smaller in its own daily use. Scores for OpenAI's models in its comparison tables were taken from OpenAI's reports, not rerun by Anthropic. Anthropic also noted that safeguard interventions may have lowered some of the model's own scores.

On safety, Anthropic rates the model's biology and cybersecurity abilities as close to its most capable models, so it carries the same safeguards as Fable 5.1. TechCrunch reported that these limit uses such as finding exploits in compiled programs. Anthropic says that when a safeguard triggers, some tasks are passed to another model, and ordinary bug fixing in software development still works. Outside groups, including METR, tested the model before release.

What it means in Nepal

The sources do not mention Nepal. The general trend they show is that capable coding models are getting cheaper and faster with each release. For any small software team or student group, lower prices make it more practical to use such tools in real projects, though the cost still depends heavily on how many tokens a task uses and which effort level is chosen. Estimating that cost before starting is now part of planning a project.

The kind of work these benchmarks measure also hints at changing expectations for engineers. The tests involve long tasks across a whole codebase: reading files, running commands, fixing errors and checking results. When a model can do more of that work, the human role shifts towards defining the task clearly, reviewing the changes, and testing whether the result is correct. Those skills come from understanding the code, not from typing it faster.

Students should also read the claims critically. Almost every number in this story comes from Anthropic or from customers it chose to quote. Independent tests usually appear weeks after a launch and sometimes tell a different story. Comparing a vendor's figures with independent results, and noting the margin of error, is a good habit for anyone choosing tools for real work.

What to study if this interests you

Software Engineering, ENCT 352, in the sixth semester of BCT, covers requirements, design, testing and code review, the skills needed to check machine-written code before it reaches users. Data Structure and Algorithm, ENCT 252, in the fourth semester, builds the ability to read and judge code quickly, which matters when reviewing large changes.

Artificial Intelligence, ENCT 351, in the sixth semester of BCT, explains how intelligent agents plan and act, and how machine learning systems are built and evaluated. The course has a full guide on this site. The Minor Project, ENCT 354, in the same semester, is a good place to practise using a coding assistant carefully: writing a clear task, reviewing every change and keeping a test suite that proves the project still works.

Words in this story

Benchmark
A fixed set of test tasks used to compare how well different systems perform.
Effort level
A setting that controls how much a model reasons before answering, trading cost and speed against quality.
Code migration
Moving a large program to a new language, library or platform while keeping its behaviour the same.
Margin of error
The range within which a measured score could vary by chance, so small differences inside it may not be real.

Where this comes from

Written in our own words; no sentence is copied from these reports. Researched with AI assistance on 11 October 2026; no member of faculty has reviewed it yet. If you spot a mistake, call 01-5091616 and we will correct it and say so.

Next story: 408IonQ says one ordinary CPU can decode quantum errors in real time

Last reviewed by Imperial College of Engineering. Written 11 October 2026 from the sources above.