From “stacking GPUs” to “squeezing tokens”: Xunce Technology (03317) TokenCloud card slot AI infrastructure new paradigm

Zhitongcaijing · 2d ago

As demand for AI computing power continues to grow and supply chain supply tightens, the AI industry is shifting from large-scale “GPU piles” to refined “token squeezing.” Tokens are also gradually evolving from simple technical units of measurement to smart assets that can be circulated and traded.

Against the backdrop of rising computing power prices and lengthening hardware delivery cycles, enterprises generally face multiple challenges such as model selection, hardware matching, scenario adaptation, and data security in promoting large-scale AI implementation.

On September 3, Xunce Technology (03317) followed the trend and launched TokenCloud's one-stop AI model training and calculation power platform to open up the value chain from data to token through software and hardware collaboration.

Breaking the bottleneck: software and hardware collaboration to improve AI input-output ratio

The Zhitong Finance App learned that TokenCloud is positioned as hardware resources and infrastructure for the transformation of data resources into tokens, and is committed to opening up a full link process from data access, computing power scheduling, model inference optimization to refining and deploying small enterprise models.

1000019543.png

In response to the problem that heterogeneous hardware such as GPUs, NPUs, FPGAs, and CPUs is difficult to unify scheduling, TokenCloud builds a unified heterogeneous computing power pooling layer, abstracts different chip resources into unified computing power units, enables unified management and intelligent scheduling across architectures, and revitalizes idle computing power resources.

Today, companies generally have the contradiction of “the big model works well but doesn't work; the small model runs but doesn't work well”. TokenCloud migrates the essence of large model knowledge to lightweight models through model evaluation and scenario optimization distillation, reducing inference delays and hardware costs while maintaining high accuracy.

At the same time, through full-link acceleration, the platform improves initial response speed and concurrent carrying capacity, so that AI applications can maintain stable operation in high-load scenarios.

In response to the most sensitive data sovereignty and hidden risk of data leakage, TokenCloud uses a cloud-side collaborative architecture to place the first and last layers of the model locally to ensure that the original data does not leave the domain, and the middle layer is deployed in the cloud to enjoy powerful computing power, achieving a balance between security and computing power efficiency, and improving customer investment return on the three dimensions of cost reduction, efficiency improvement, and accelerated AI implementation.

Build a closed ecological loop and open up new space with domestic computing power

From an industrial chain perspective, TokenCloud is not an isolated hardware scheduling tool, but a key puzzle that is deeply linked with Xunce's underlying AIDP data resource processing system, AI and token operating system TokenOS, and the higher-level model result service and scenario connecting layer TokenRouters.

1000019545.png

Among them, TokenOS focuses on solving algorithm selection and software engineering problems from data to token generation, while TokenCloud is responsible for heterogeneous computing power adaptation and hardware infrastructure support. The two work together to achieve joint decisions on model optimization and computing power optimization, so that each computing power can accurately match the business scenario.

The Zhitong Finance App learned that the deep moat built by Xunce in the AI implementation chain stems from its proprietary data understanding and customer trust accumulated over the past ten years in high-threshold, high-viscosity vertical industries. This unique data set and engineering experience based on specific production and operation scenarios has rapidly reduced the team's marginal service costs in vertical scenarios.

At the same time, the company relies on an FDE model close to the front line of business to quickly adapt standardized product capabilities to complex business bottlenecks, effectively resolving the conflict between individual requirements and standardized products. Currently, the company's business covers 11 high-value industries such as finance, telecommunications, electricity, high-end manufacturing, and biomedicine. It has successfully achieved commercial verification in core scenarios such as biomedicine and industrial quality inspection, and has an in-depth understanding of customer needs and pain points.

Overall, in the past, the company solved more about how to convert raw data into high-quality data; now it is expanding further upward to solve how to form professional models and industry scenario tokens based on high-quality data, and further improve the circulation efficiency and commercial value of tokens. This also forms a closed loop of the company's complete product from computing power, data, tokens, models to applications.

Based on a broader industrial and capital perspective, TokenCloud has demonstrated significant long-term strategic value and ecological card advantages:

First, the platform adapts to the business paradigm shift from “selling cards” to “selling tokens” in the computing power industry, enabling the company to take an advantageous position under the new economic paradigm with Token as the core billing unit;

Second, in the critical period of AI full-stack localization entering the system engineering optimization stage, TokenCloud has established deep cooperation with many domestic GPU manufacturers, becoming a key hub connecting domestic heterogeneous computing power with enterprise-level AI applications, and is in deep line with the policy orientation of the construction of a national integrated computing power network and high-quality development of computing power infrastructure.

According to the opinions of many mainstream brokerage firms and institutions, under the trend of AI agents driving a nonlinear explosion in demand for inference computing power, infrastructure vendors with full-link engineering implementation capabilities and the ability to convert electricity and chips into usable token computing power at a lower cost will be the first to enter the dividend cycle of accelerated release of performance.

In this process, Xunce Technology uses TokenCloud as a fulcrum to further cultivate business scenarios and connect algorithms and computing power ecosystems upward, and is facing a strategic reevaluation opportunity to leap from a “digital base” to an “AI productivity platform.”