Citi: Improving AI Reasoning Efficiency, Dramatically Reducing Task Costs with Frontier Proprietary Models

Zhitongcaijing · 1d ago

The Zhitong Finance App learned that Citi released a research report saying that the latest cutting-edge model has shown improvements in performance and cost of each task, highlighting improvements in model efficiency and economic efficiency of reasoning, and may help strengthen the long-term economic advantages of relatively less efficient models and open source weighted alternatives. The bank pointed out that the top proprietary model released this week increased Artificial Analysis (AA) intelligence frontier by 9% and reduced AA task costs by 55% compared to its predecessor model; the release of the cutting-edge model accelerated, continuing to re-widen the gap with the open source model.

According to the bank, the cost of each task in the open source model fell 35% to 0.8 US dollars on a weekly basis, and the discount compared to the closed model increased from 40% to 60%; the intelligent lead of the proprietary model increased from 9 points to 12 points. Cache-related operations rose to 61% of task costs, showing that the next stage of inference efficiency may come from providing more context at a lower marginal cost, benefiting vendor margins, and making more proxy workloads possible. The US-EU hybrid token price fell 25% from week to week, but the Silicon Data token spending index rose 3%, reflecting the elasticity of demand at lower prices. The number of proprietary models released reached 20 last week, which is about 67% higher than the 12 in the past five weeks.

The bank also pointed out that the MiMO-v2.6-Pro of Xiaomi Group-W (01810) shows that the cost of obtaining capacity improvements through agentic RL has become low: operating with an RL of about 2.6 million US dollars for a period of 6 days on an existing basic model, a top-level open source weighting model is produced. It is expected that cutting-edge laboratories will increasingly replicate similar processes in their own model series. Consumer proxy apps quickly became popular. Grok Bot recorded 935,000 iOS downloads within 44 days, and Muse surpassed 2 million in 15 days; however, computing power limitations and the zero-sum reality of the token economy may allow suppliers to balance consumer adoption with higher-value enterprises and wholesale revenue. Personalized, always-on autonomous agents may drive up CPU and network requirements in the short term. As agents accumulate context, become more personalized, and take on longer tasks, long-term GPU demand may also increase. This Tuesday (29th), OpenAI DevDay is expected to provide more information on personalized agency architectures and their business acceleration.