Business

India PR Distribution
New Delhi [India], October 8: For most of the past two years, the frontier of artificial intelligence has had a simple price tag. If you wanted the strongest models, you paid the strongest rates, and the strongest rates belonged to a small club of well-funded American labs.
--By Wei Wang
Ethos 6 is testing whether that rule still holds.
The model, developed by Atomesus AI, has just been put through an extensive third-party benchmark covering more than two dozen categories, from software engineering and debugging to biology, cybersecurity and statistics. The results, which were checked against the underlying result files rather than taken from a promotional chart, place Ethos 6 well above where a budget model is supposed to land. In the broadest of the three test suites, it averaged 92.68 out of 100.
That would be notable on its own. What makes it hard to ignore is the price. Through its hosted API, Ethos 6 now costs $0.52 per million input tokens and $2.07 per million output tokens, according to the company's pricing page. OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1, the two models most often cited as the current frontier, are listed at $10 and $50 for the same amounts.
Put plainly, a workload of one million tokens in and one million tokens out costs $2.59 on Ethos 6. On either of its best-known rivals, it costs about $60. That is not a discount. It is a different business model, and it raises a question the industry has mostly avoided: how much of the frontier premium is paying for capability, and how much is paying for the brand?
This article reviews the third-party benchmark results for Ethos 6 alongside the published figures for Astra and Fable 5.1. The author has no connection to Atomesus AI. One point needs to be clear from the start, and it will be repeated throughout: the Ethos 6 scores come from a third-party benchmark package that uses different tests from the ones OpenAI and Anthropic report. They describe the shape of the model's abilities. They do not settle which model is best.
The two models everyone measures against
To understand why the Ethos 6 results drew attention, it helps to start with the models it is being compared to.
OpenAI's GPT-6 Astra is the company's million-token flagship. It reads up to 1.05 million tokens at a time, writes up to 128,000 tokens in a single response and lets developers set its reasoning effort anywhere from low to maximum. OpenAI markets it for reasoning, coding, research and, above all, controlling a computer on a user's behalf.
Anthropic's Claude Fable 5.1 matches Astra almost spec for spec. It holds a one-million-token context window, produces up to 128,000 tokens of output and runs adaptive thinking at all times. Anthropic describes it as the most capable model it offers to all customers and pitches it for agents that work on their own for hours.
On the tests the two companies share, the race is close. In OpenAI's own head-to-head comparison, Astra scored 97.6 percent on FrontierMath Tier 4 against Fable's 87.8 percent, and 96.0 percent on GPQA Diamond against 93.7 percent. On the DeepSWE v1.1 coding benchmark, Astra led 74.1 to 67.4. On Terminal-Bench 4.0, the margin was just 57.9 to 55.8.
Fable won where the work was broader. It scored 65.0 percent on Humanity's Last Exam with tools, against Astra's 57.2 percent, and led the Artificial Analysis Intelligence Index at 65.7 to 61.2.
The picture that emerges is not a single winner but a division of strengths. Astra leads in hard mathematics, coding and computer control. Fable leads in long, research-style reasoning that draws on large amounts of context. Into that picture steps a third model that does not appear on either company's charts at all.
What the third-party benchmark found
The third-party evaluation of Ethos 6 is not a single test. It is three separate suites, each with its own mix of categories. Read in order, they tell the story of a model that looks less perfect, and more believable, the harder it is pushed.
The first suite produced a wall of perfect scores. Ethos 6 reached 100 in 14 of 29 categories, including mathematics, debugging, cybersecurity, visual reasoning and computer-aided design. Then the wall cracked. Tool use came in at 67.5, long-context reasoning at 57.5, and data interpretation, computer systems and graph reasoning sat at or near 50. Two figures in that suite, 53 for factual omniscience and 60 for repository-level software engineering, were supplied by the publisher after the original result files were produced. They are revisions rather than reproduced results and should be read with that in mind.
The second suite was more balanced and averaged 89.63 out of 100. Language understanding, algorithms, mathematics and debugging all scored a perfect 100. Knowledge and factual analysis reached 99.3, cybersecurity 99.2 and biology 98.8. Here, too, the weaker areas were clear: agentic planning fell to 58.5, economics and finance to 72, and statistics to 78.5.
The third suite is the broadest of the three and, for most readers, the most persuasive. Across 20 categories, Ethos 6 averaged 92.68 out of 100. Its single highest score was not in mathematics or coding but in writing and editing, at 98.52. Language understanding followed at 98.40 and software engineering at 98.32. Databases scored 97.44, multistep problem solving 97.40 and debugging 96.92. Cybersecurity and planning and optimization each reached 96.00, and factual analysis reached 95.00.
The middle of the range was still strong. Statistics and probability scored 94.84, mathematics 94.32, logic and reasoning 94.20 and data interpretation 93.20. In the sciences, chemistry came in at 92.72, biology at 91.08 and physics at 90.20. Algorithms scored 88.40 and computer systems 84.16.
The pattern in that third-party data is worth noticing. The best results cluster in tasks where information has to be organized, transformed, repaired or carried through many dependent steps. That is the profile of a capable professional assistant rather than a quiz champion. It suggests a model tuned for the kind of everyday knowledge work that most paying customers actually do.
The weak spots, and why they matter
No model is good at everything, and the third-party results do not pretend otherwise.
The most practical weakness is instruction following, which scored 78.40 in the third suite. That is 20 points below the model's score for language understanding. In everyday use, a gap like that tends to show up in a specific and frustrating way: the model understands exactly what you asked for, then quietly ignores one of your rules. It might write a strong summary but run past the word limit, or produce clean code in a style you told it not to use. For businesses building products on top of the model, that is the kind of behavior that has to be caught and corrected in testing.
Economics and finance was another soft spot. It finished last in the third suite at 78.16, and it was also among the weakest categories in the second suite, at 72. Anyone planning to use Ethos 6 for financial analysis would be wise to check its work closely.
The agent-related numbers are the most mixed. Planning and optimization scored a strong 96.00 in the third suite. But agentic planning, measured in the second suite, came in at just 58.5, and tool use in the first suite was 67.5. Those are the abilities that matter most for AI agents that work on their own over long tasks, which is exactly the ground where Astra and Fable are competing hardest.
Oddly, this unevenness is one of the more reassuring parts of the record. A model that scores 100 on everything usually reveals a test that is too easy, not a model that is too good. The fact that the third-party benchmark found real gaps, some of them large, makes the high scores elsewhere easier to believe.
Who held the stopwatch?
In any story about AI benchmarks, the most important question is not the score. It is who ran the test, and how.
The Ethos 6 figures come from a third-party benchmark package, not from Atomesus AI's own marketing. That matters. Company-run benchmarks are often chosen, tuned and presented in whatever way flatters the product. A third-party benchmark removes some of that risk, and in this case the scores were checked against the raw result files.
But third-party does not mean identical. The Ethos 6 tests were not run on the same harness, with the same settings, as the tests OpenAI and Anthropic publish for their models. That distinction is easy to blur, so it is worth spelling out. Ethos 6's score of 98.32 in software engineering does not mean it beats Astra's 74.1 percent on DeepSWE. Astra's 97.6 percent on FrontierMath does not prove it is better at mathematics than Ethos 6, which scored 94.32 in that category. These are different exams with different tasks and different scoring rules. Comparing the raw numbers is like comparing a marathon time with a high-jump record.
What can fairly be compared is the shape of the evidence. Astra's published record is deepest in mathematics and computer control. Fable's is deepest in long, context-heavy research. The third-party record for Ethos 6 shows a model that scores in the 90s across most of the categories that professionals care about, with clear weaknesses in instruction following, finance and some agent tasks.
By that measure, Ethos 6 looks much more like a frontier contender than a cheap lightweight. Whether it truly belongs alongside Astra and Fable is a question only a shared, controlled test can answer.
The race to build AI that works on its own
The next big contest in AI is not chat. It is agents: models that take control of a computer, use tools and keep working on a task for hours without a person at the keyboard.
OpenAI has made this Astra's calling card. The company reports a score of 59.3 percent on Agents' Last Exam, 72.6 percent on the August 2026 OSWorld 2.0 offline set under its partial-scoring setup, and 92.7 percent on ScreenSpot-Pro without tools. It says Astra finished its OSWorld run in roughly 47 percent less time than its predecessor, GPT-5.6 Sol. On AutomationBench, Astra scored 41.4 percent against Fable's 31.4, and on BenchCAD it led 95.9 to 84.3.
Anthropic tells a different story about Fable 5.1, one centered on endurance. In its own evaluation, the model scored 52.6 percent on Terminal-Bench-Science 0.1, 55.8 percent on Terminal-Bench 4.0 and 73.4 percent on CursorBench 3.2.0. On OSWorld 2.0, Anthropic reports 77.9 percent under partial scoring and 41.7 percent under strict scoring. Because the two companies used different task releases and setups, Fable's partial score cannot be placed directly against Astra's 72.6.
Ethos 6 wants a seat at this table, and the third-party data gives a mixed answer. Its 96.00 in planning and optimization and 97.40 in multistep problem solving point to a model that can reason through long chains of work. Its 58.5 in agentic planning and 67.5 in tool use, from the other two suites, suggest it is not yet as dependable when it has to act on its own.
That contradiction is not unusual. How well an agent performs depends heavily on the scaffolding around the model: the tools it can call, the rules for retrying a failed step, how it is given feedback and how its memory is managed. A good harness can lift a model's agent scores sharply, and a poor one can sink them. Until Ethos 6 is tested on the same agent benchmarks as its rivals, its real standing in this race is an open question.
The benchmark that pays the bills
For developers and companies, this is where the debate stops being academic.
Ethos 6 is offered as a hosted model API under the name ethos-6. According to the company's pricing page, it costs $0.52 per million input tokens and $2.07 per million output tokens. Billing is in US dollars and is drawn from a prepaid wallet, so customers top up a balance in advance rather than receiving a bill at the end of the month. The company's own example puts one million input tokens plus one million output tokens at $2.59.
Astra and Fable 5.1 are both listed at $10 per million input tokens and $50 per million output tokens. That makes Ethos 6's input about 95 percent cheaper and its output about 96 percent cheaper. The same million-in, million-out workload that costs $2.59 on Ethos 6 costs about $60 on either rival, roughly 23 times as much.
The gap grows with scale, and output is where it hurts most. Ten million output tokens cost about $20.70 on Ethos 6 and $500 on Astra or Fable. At one hundred million output tokens, the comparison is $207 against $5,000. For a company running an AI feature for thousands of users, that difference can decide whether the feature makes money or loses it.
There are fair caveats. Both rivals offer cached-input pricing that can sharply cut costs for workloads that reuse the same long prompt. OpenAI lists Astra's cached input at $1 per million tokens, and Anthropic lists cache reads for Fable at $0.25. Ethos 6's pricing page does not list a cached rate, so for heavily cached workloads the real-world gap may be smaller than the headline numbers suggest.
Ethos 6 also uses an OpenAI-compatible API. That is a practical detail with large consequences. A developer who has already built a product on OpenAI's interface can, in many cases, try Ethos 6 by changing an address and a model name rather than rewriting code. It lowers the cost of experimenting, which is exactly what a challenger needs.
To pull in new customers, the company is also running a first top-up offer. Users who add $10 or more to their wallet receive a one-time bonus worth the Indian rupee equivalent of $100 in credit. The bonus is added only after a successful, verified payment and is available once per account. The rupee wording, along with the oddly precise per-token prices, hints that the company may set its prices in rupees and convert them to dollars, which could mean the listed dollar rates shift slightly with exchange rates.
A $1.25 plan in a $20 world
The consumer pricing is just as aggressive. Ethos 6 is sold in three monthly plans, each of which, the company says, comes with a promise that user data is never used for training.
The entry plan, Core, costs $1.25 a month and is pitched as smarter AI for everyday work. It includes access to Ethos 6, higher messaging limits and more file uploads and image generations.
Prime costs $4 a month. On top of everything in Core, it adds what the company calls advanced reasoning and problem solving, significantly higher usage limits, more uploads and image generations, and an extended context window with longer memory.
The top tier, Plus, is listed at $20 a month but is currently offered at $10, a 50 percent discount. It adds unlimited Core chat, the highest usage limits, unlimited file uploads and image generation, the longest context window and maximum memory, expert reasoning and advanced coding, and deep research with priority processing.
For comparison, $20 a month has become the standard price for a premium AI assistant subscription. Ethos 6's top plan, at its current promotional price, costs half that, and its entry plan costs less than a cup of coffee.
How much those plans are really worth depends on details the pricing page does not spell out. Terms like higher limits and significantly higher usage limits are not numbers, and a cheap plan with tight caps is not the same as unlimited access to a frontier model. The discount on Plus also has no stated end date. Readers weighing a subscription should look for the actual message caps before committing.
Still, the strategy is plain. Ethos 6 is not trying to be another $20-a-month assistant. It is trying to make price itself the headline.
So which one should you choose?
The honest answer depends on the job.
For demanding computer control, frontier-level mathematics, automation and heavily tested coding work, Astra has the deepest body of directly published evidence among the three. Its results on FrontierMath, GPQA Diamond, DeepSWE, AutomationBench and computer-use tests make the strongest standardized case.
For long research sessions, sustained coding, context-heavy agents and workflows that reuse a large cached prompt, Fable 5.1 is built for the purpose. Its million-token context, always-on adaptive thinking and low-cost cache reads are designed for exactly that kind of work.
For everything else, and especially for anyone whose budget is the deciding factor, Ethos 6 deserves a serious look. The third-party benchmark does not show that it is smarter than Astra or Fable, and no fair reading of the data claims it does. What it shows is that a model priced at a small fraction of the frontier now scores above 96 in writing, language understanding, software engineering, databases, multistep problem solving, debugging and cybersecurity. Models in the budget tier have not traditionally looked like that.
If those results hold up in real-world use, $0.52 and $2.07 per million tokens becomes one of the most important numbers in the industry. That if still carries real weight.
What Ethos 6 needs next is not another high score on its own terms. It needs a controlled run on the same tests its rivals already face: Terminal-Bench 4.0, FrontierMath, GPQA Diamond, Humanity's Last Exam, DeepSWE, AutomationBench, OSWorld and a recognized independent intelligence index, all with the same rules and disclosed settings. A third-party benchmark has opened the door. A shared one would settle the argument.
Until then, the fair verdict is this. The three models are not interchangeable, they do not share a scoreboard and none of them can be crowned best at everything.
But not long ago, a model priced like Ethos 6 would never have been invited into this conversation. Now, on the strength of a third-party benchmark, it has been. That may be the biggest story of the three.
(ADVERTORIAL DISCLAIMER: The above press release has been provided by India PR Distribution. ANI will not be responsible in any way for the content of the same)