Cerebras Systems (CBRS.US) Q2 Conference Call: Cloud Revenue Surges 287% to Benefit from OpenAI's Volume, AWS Bedrock Launches Q1 Next Year to Open a Hyperscale Blue Ocean

Zhitongcaijing · 2d ago

Zhitong Finance App learned that after the US stock market on Wednesday, AI chip star Cerebras Systems (CBRS.US) announced record second-quarter core revenue and raised its full-year results guidance. CEO Andrew Feldman said that all of the company's core indicators have exceeded expectations. Core revenue is expected to more than double in 2027 and maintain multiple growth over the next few years. Management has positioned 2026 as the “groundbreaking year” to prepare for more than $25 billion in remaining performance obligations (RPO).

Specifically, core revenue for the second quarter reached US$209.9 million, an increase of 103% over the previous year. Among them, core cloud and other service revenue was US$127.7 million, a year-on-year increase of 287%; core hardware revenue was US$82.1 million, an increase of 17%. The core gross margin was 40.6%, up about 940 basis points from the same period last year, but there was a month-on-month decline from 46.5% in the first quarter, mainly because the company temporarily leased back some systems from cloud customers to meet private cloud needs. This high-cost leasing capacity reduced gross margin by about 500 basis points. The company expects a low gross margin in the third quarter, which will improve with the launch of its own system data center in the fourth quarter, with a long-term target of over 60%. Core operating losses were $33.6 million, and operating margin improved to -16% from -42% in the same period last year.

Looking ahead to the third quarter, the company forecasts core revenue of $214 million to $216 million, gross profit margin of 38% to 40%, and operating margin of -25% to -23%. Core revenue expectations for the full year were raised to US$880 million to US$890 million, with gross margin guidance of 41% to 43%, and operating margin of -19% to -17%.

In terms of capacity expansion, data center space is still a bottleneck in the industry. The company has locked in more than 600 megawatts of capacity (operation or contract until the end of 2027). Future expansion reserves are measured in gigawatts, and the layout covers many states of the United States, France, Finland, Canada, etc. The manufacturing side expanded production through Flex and Sanmina, and production capacity increased fourfold compared to the first half of 2025. It is expected to exceed tenfold in 2026, and then increase three to four times in 2027. The supply of 5nm wafers with TSM.US is stable, and there is no need for high-bandwidth memory, CoWOS packaging, or 3nm production capacity, reducing supply chain risks.

Technically, support for the OpenAI GPT-5.6 Sol model was added this quarter, making it ten times faster. The company promotes decoupled inference cooperation with AMD (AMD.US) and Amazon (AMZN.US) AWS. The GPU processes the pre-filling stage, and the Cerebras system is responsible for decoding. The goal is to increase throughput by five times. It is expected to be deployed in the fourth quarter. The company will release the fourth-generation CS-4 system at the Supernova conference. CS-5 is scheduled to be launched in the second half of 2027. The performance will double every year for the next few years, and the throughput will increase by more than 20 times within 18 months.

In terms of customer development, Cerebras expects the product to be fully launched through the AWS Bedrock platform in the first quarter of 2027, generating the first hyperscale revenue in the middle of the year, but the RPO of 25.4 billion US dollars as of June 30 does not include hyperscale vendors such as AWS. Six more than $30 million deals were signed in the second quarter. Customers include Figma (FIG.US), Cognition, Lovable, Block (XYZ.US), AlphaSense, GlaxoSmithKline (GSK.US), and CrowdStrike (CRWD.US). OpenAI will still account for an important share next year, but AWS, coding, and security applications will gradually expand, and “new cloud” vendors will become an important part of their business in 2027.

Below is the full text of the Cerebras Systems results conference call

operator

Good afternoon, and welcome to the Cerebras Systems Q2 FY2026 earnings conference call. Please note that today's call will be recorded. I'll now forward the call to Sean Dorsey, Head of Investor Relations. Please get started.

Sean Dorsey

Thanks, operator. Good afternoon, and welcome to the Cerebras Systems FY2026 second quarter earnings conference call. Earlier today, we issued a press release and a supplementary earnings presentation in the investor relations section of our website. A replay of this webcast will also be made available on our investor relations website after the conference. Joining me today are Andrew Feldman, our co-founder, CEO and president, and Bob Komin, our chief financial officer.

Before we begin, I would like to remind you that today's discussion will include forward-looking statements under the safe harbor provisions of the Private Securities Litigation Reform Act of 1995. These statements include, but are not limited to, statements about our future financial performance, business strategy, market opportunities, customer needs, product roadmap, technology leadership, supply chain, operating model, and outlook for the third quarter and full year of 2026.

Forward-looking statements are based on our current expectations and assumptions, and are subject to various risks and uncertainties that may cause actual results to differ materially from those expressed or implied in the statements. These risks are described in our SEC filings, including the final prospectus relating to our initial public offering and our future regular filings with the SEC.

We are under no obligation to update these forward-looking statements unless required by law. In today's conference call, we'll also be discussing some non-GAAP financial measures. The reconciliation between GAAP and non-GAAP results is included in today's press release and supplementary materials, which are available on our website's investor relations page. Well, I'll give Andrew the call.

Andrew Feldman

Co-founder, CEO, President and Chairman

Thanks, Sean. Thank you all for attending our conference today. The second quarter was strong. We completed our public offering, but that didn't distract us from carrying out our established plans. We achieved record core revenue and surpassed expectations on all metrics: core revenue, core gross margin, and core operating margin.

Looking ahead, we see that the need for fast inference (fast inference) is endless. The market is realising that speed isn't just an indicator for benchmarking. Speed has changed user engagement, changed the performance of agents (agentic), and changed the productivity of artificial intelligence. Quick inference unlocks new apps and new markets.

As we've shared with you before, 2026 is a groundbreaking year for Cerebras. In the seven weeks since the last quarter earnings call, we have made outstanding progress in many areas, preparing us for significant developments in 2027, 2028, and 2029 to meet the $25 billion remaining performance obligation (RPO) we currently have on our books.

Thanks to these developments, we expect core revenue to more than double in 2027 and continue to grow several times over the next few years. We believe progress is reflected in three areas: production capacity, capacity, and customers.

We're expanding production capacity by adding new data center contracts, expanding manufacturing capacity, and partnering with suppliers around the world to ensure supply and support our extraordinary growth. We're improving our capabilities by inventing new technologies that extend our strengths in performance, throughput, and energy efficiency. We're expanding our customer base by accelerating AI productivity in existing markets such as code generation and agent workflows, and opening up new areas such as security — in these new areas where speed is opening up unprecedented opportunities.

In terms of production capacity, data center space remains a bottleneck for the entire industry, and we are no exception. The sooner we and our customers can launch new data centers, the faster we will grow. Therefore, over the past seven months, we have worked hard to acquire and build a data center. We have two advantages. First, since we serve the inference business, we don't need the kind of gigawatt (gigawatt) infrastructure needed to train a cluster. This gives us more flexibility when expanding our production capacity at multiple locations around the world.

Second, we've established a repeatable process for site selection, cluster deployment, and customer activation, which is the operational capability necessary to turn gigawatts of electricity into usable production tokens (production tokens) on a global scale. I am happy to report that our efforts have been very successful. We now have operating or contracted data centers in Alabama, Dallas, Denver, Minneapolis, Santa Clara, Stockton, and locations outside the US in France, Finland, Manitoba, Montreal, Norway, Saskatchewan, and Toronto.

Overall, over the past seven months, we've secured over 600 megawatts of data center capacity, which is either online or will be delivered before the end of 2027. Although this is far from meeting our needs, our data center project reserves (pipelines) for future expansion continue to grow and currently reach the gigawatt scale. Looking at it another way, as we continue to build a first-party cloud (first-party cloud), it will become one of the largest artificial intelligence clouds other than hyperscale cloud service providers.

At the end of 2025, we were still going through a steep learning curve. Today, I'm happy to report that we are already quite skilled in data center construction and have a clear path to excellence. Another key dimension of production capacity is manufacturing and supply chain. In this regard, we have successfully increased our manufacturing capacity and partnered with Flextronics and Sanmina to build a new plant. We expect to increase our manufacturing capacity by more than tenfold in 2026 and continue to expand this capacity in 2027 — again, to prepare for the extraordinary growth expected in the next few years.

Our partnerships with supply chain suppliers have also turned into significant advantages. Once again, TSMC has given us strong support, and we have the wafer supply we need to drive growth. Our ability to obtain a supply of wafers is also due to the fact that we are able to achieve industry-leading performance on TSMC's 5nm node, which has a lower wafer cost and relatively less tight supply.

Our decades-long relationships with supply partners have further strengthened our confidence in achieving future growth plans. These relationships are scarce and valuable, especially during times of tight supply. Finally, keep in mind that most of the key supply chain restrictions the industry is currently facing don't apply to us. For example, we don't use HBM memory, CoWoS packages, and we don't need a wafer manufacturing capacity of 3 nm.

In terms of capabilities, in the second quarter, we supported OpenAI's GPT-5.6 Sol, which is the largest and most capable cutting-edge model. In fact, Cerebras provides GPT-5.6 Sol services 10 times faster. With GPT-5.6 Sol, any lingering doubts about our ability to support large cutting-edge models have been dissipated. Being a delivery partner for GPT-5.6 Sol and servicing it on our cloud fully illustrates the maturity of our software stack.

Achieving a level that can provide large-scale quality and reliability requires millions of hours of system testing in harsh production environments. We're proud that our inference cloud can meet the requirements of our most demanding customers. Our collaboration in providing services for cutting-edge models has opened up new and important strategic advantages that were previously only available to NVIDIA (NVIDIA). Frontier closed-source models include a continuous stream of new insights and technology.

Providing services to these models allows us to see and prepare for the future. Our hardware and software stack roadmap now reflects the trends we're seeing and will provide us with compound benefits for years to come. Continuing around the topic of “competencies,” let's talk about “disaggregation” (disaggregation). We now have decoupled inference solutions with two leading chip companies — AMD (using its Helios system) and AWS (using its Trainium chip).

Decoupling has expanded the market for GPU vendors and Cerebras. Decoupling allows GPUs to participate in a market that is currently off-limits to them, the fast reasoning market. The decoupling enabled Cerebras to expand our opportunities to those more price-sensitive customers and increase the profitability of our data centers. Let's see how it works. As with any computing market, as reasoning grows and matures, opportunities for specialization arise.

Decoupling is a form of specialization, and is particularly suitable for workloads with known traffic patterns. In these cases, decoupling provides an advantage by dividing inference into two stages — pre-filling and decoding, and using a different processor for each stage. Prefill processes input from users or agents. This is a parallelizable workload. As a result, pre-filling is ideal for GPUs and their HBM-based memory architectures. Decoding is responsible for generating output tokens.

This is a more difficult technical problem, and is where most of the computational work in the decoupling solution lies. It is sequential and has extremely high memory bandwidth requirements, making it a perfect fit for our wafer-level engines. Pre-populated and decoded processors need to be connected to form an end-to-end solution. And in this regard, our standards-based I/O and open collaboration strategy make Cerebras integration simple and straightforward.

A few weeks ago, we announced our partnership with AMD to build a decoupled inference solution. These solutions combine their Helios rack with our CS system. The combined solution increased throughput by 5 times while maintaining Cerebras speed. The key to understanding how powerful this is is to differentiate between “speed” and “throughput.” Speed is a measure of an individual user.

It is measured in “tokens per second per user”. It measures how quickly your query is responded to, or how long it takes an agent to complete a task. This is the content on the X axis. Throughput, on the other hand, is the total number of tokens the solution can generate per second. It is measured by adding up all tokens for all concurrent users. Here, it's on the Y axis as usually shown. Speed is critical to the user experience. Throughput is critical to the economic efficiency of inference.

GPU solutions can support high throughput, but only at low speeds. When configured to support even moderate speeds, the GPU's throughput drops dramatically. This applies not only to GPUs, but also to ASICs and all solutions using HBM. The HBM memory architecture forces a trade-off between throughput and speed. SRAM-based architectures like Cerebras are the complete opposite. We support extremely fast token speeds, but the throughput is moderate. As a result, GPUs want to be faster without sacrificing throughput.

Cerebras wanted higher throughput without sacrificing speed. This is the strength of our decoupled solution. It provides the speed of Cerebras with 5x higher throughput. While maintaining our industry-leading speed, increasing throughput by 5 times has had a profound impact on the economic efficiency of each token generated. This means that each Cerebras system can generate up to 5 times faster, high-value tokens. Each system produces more tokens and costs less, which means higher revenue and higher gross profit margins.

Generating more tokens per CS system also means more tokens per watt, making every data center more profitable. Perhaps most importantly, in an environment where data center resources are limited, decoupling solutions enable us to meet more of our RPO requirements. Finally, we believe this decoupling method can provide performance and economic benefits on any GPU.

For operators that have already deployed a large number of GPUs, decoupling with Cerebras provides them with an opportunity to significantly increase the value and utility of their data center assets by pairing some of their GPUs with Cerebras solutions to create significant leverage on their existing investments. Continuing on the theme of “Competence,” let's move on to our roadmap. Our project execution is progressing steadily.

We expect to double the speed of delivery of new systems every year over the next few years. Remember, we're doubling our performance on top of already having a performance advantage that's 15 times faster than anyone else in the industry. Additionally, while maintaining leading performance over the next 18 months, we plan to deliver a solution that will increase throughput by more than 20 times. Next week, at our annual Supernova conference, we'll unveil our fourth generation system — CS4.

It's going to be a big event with lots of product launches, so I recommend everyone to attend. Finally, we're currently on track to launch our CS5 in the second half of 2027. Looking further into the future, our innovation engine is roaring. We have important partnerships with the US government in delivering stacked memory solutions as well as integrated wafer-level optical solutions.

Over the next few years, you can expect us to innovate in chips and chip architectures, as well as in all aspects of system design, including packaging, I/O, and power delivery. Summarizing the “Capabilities” section: We look forward to continuing to deliver groundbreaking advancements in products and technology to increase speed and throughput, reduce energy consumption per token, and drastically reduce the cost per token of our solution.

Now let's turn to the customer side. Fast tokens are in high demand and enjoy premium prices in the market, while fast tokens with cutting-edge intelligence are only available through the partnership between OpenAI and Cerebras. Our collaboration with AWS continues, and we expect our solution to be officially launched via AWS's Bedrock platform in the first quarter of 2027. This AWS partnership has expanded our market opportunities and provided us with a global reach through an industry leader trusted by almost all businesses around the world.

Our discussions with other hyperscale cloud service providers are also progressing well. We expect to generate our first revenue from mid-2027 and gradually scale up in 2028 and beyond. While making all of this progress, I think it's important to remember that our $25 billion RPO currently doesn't reflect any backlog from AWS or any other hyperscale cloud service provider.

Our business outside of OpenAI and hyperscale cloud service providers continues to grow healthily. For example, in the second quarter, we signed 6 deals worth over $30 million. AI code generation continues to grow rapidly. In our experience, no one is happy with slow tokens when programming. So it's no surprise that our influence in the field of code generation continues to expand. We have signed new agreements with listed companies such as Figma and startup leaders such as Cognition, and our business in Europe has also made significant progress, and we have reached an important partnership with Lovable.

Agent workflows are growing rapidly, and as agent operations rapidly evolve to multi-step, multi-agent solutions, the value of speed continues to accumulate. Several companies, including Block, AlphaSense, and GSK, signed new agreements with Cerebras in the second quarter to use fast reasoning to support their customized AI agents. Rapid AI has also opened up new markets and expanded Cerebras' potential market reach (TAM). The field of security is an example.

Our recent collaboration with CrowdStrike is an app that can only be achieved if AI is fast enough. Fast AI enables AI-based security devices to be embedded in enterprise traffic and uses big language models to secure traffic so quickly that no one is aware. Fast AI enables big language models to provide security that is invisible to users. AI provides security, and speed creates this “stealth,” enabling security to avoid delays and outages.

Given the rapidly evolving threat landscape, we expect this type of security to become the norm. Businesses will soon expect the vast majority of their traffic to be inspected in this way, which will create huge new opportunities exclusively realized by fast AI. Frontier laboratories, hyperscale cloud service providers, leading chip makers, the fastest growing startups, and large enterprises are now Cerebras customers and partners and benefit from our extremely fast inference services.

All in all, it's been a strong quarter. We successfully completed our IPO. We surpassed expectations on all metrics: core revenue, core gross margin, and core operating margin. We have made progress in three key areas — capacity, capacity, and customers. These are the cornerstones for us to achieve significant growth in 2027 and 2028 and continue this extraordinary growth rate in 2029 and beyond. So, I'll hand the phone to Bob. Bob?

Robert Komin

Senior Vice President, Chief Financial Officer and Treasurer

Thanks, Andrew, and good afternoon everyone. We made tremendous progress in the first half of 2026. As we have described, 2026 is the year that lays the foundation for several times growth over the next few years. At the beginning of the year, we won one of the largest technology deals ever, generating an RPO of over $25 billion. This requires us to immediately begin large-scale expansion in the three key components of production capacity.

First, we need to increase the supply of wafers. As Andrew described, thanks to our strong relationship with TSMC and strong support, we have achieved this and are fully prepared not only for the rest of the year, but also for next year. Second, we need to expand manufacturing capacity. Our manufacturing capacity is now 4 times what it was in the first half of 2025, and will increase it by more than 10 times in 2026. As a result, we have made tremendous progress in this area.

Third, we need a significant increase in data center capacity. We have made significant progress, with over 600 megawatts of capacity already online or contracted, and expected to be delivered by the end of 2027, in addition to gigawatts of project reserves. As a result, the foundation is in place to support core revenue growth of two or more times in 2027 and achieve additional multi-fold growth over the coming years. This growth also paved the way for a significant expansion of our profit margins in 2027 and beyond. Now let's take a look at the financial results for the second quarter.

We have ushered in another strong quarter, exceeding expectations in terms of performance guidance indicators. We have achieved record core revenue. We also surpassed expectations in terms of core gross margin and core operating margin. I'll continue to use the core business framework introduced last quarter to describe our progress. The definition of core business metrics and the reconciliation of all metrics to GAAP are included in today's earnings report and on our website.

Core revenue was $209.9 million, up 103% year over year. Our private cloud business is growing at an incredible rate. Core cloud and other services revenue was $127.7 million, up 287% year over year. This nearly four-fold increase reflects the huge market demand for our Cerebras rapid inference services. Core hardware revenue for the quarter was $82.1 million, up 17% year over year.

We're focusing on total core revenue rather than the component ratio of the two, as this ratio may fluctuate significantly from quarter to quarter due to the timing of large-scale cloud capacity additions and the timing of hardware shipments. In the second quarter, most of the increase in total core revenue was due to growth in our core cloud services, which reflected the acceleration of our OpenAI deployment, increased usage by other cloud customers, and some hardware customers also struggling with the schedule for launching new data center capacity.

Demand for quick reasoning remains strong. Currently, there are several late-stage hardware deals, with potential orders worth hundreds of millions of dollars from new customers, as well as important new cloud service deals for 2027. The existing fast reasoning market is growing, and new markets are being launched. We see decoupling as an important tool for driving new application scenarios because it can significantly improve the unit economics of inference business and data center ownership.

Currently, this means that each CS system can generate up to 5 times more tokens, so power consumption and cost per token are significantly reduced. Furthermore, through continued strong investment in R&D and product roadmaps, Cerebras will quadruple our current industry-leading speed and increase throughput by more than 20 times by the end of 2027, thereby greatly improving our performance and the unit economic efficiency of inference.

Now let's look at gross profit margin. On a year-on-year basis, core gross margin increased significantly. The core gross margin for the second quarter was 40.6%, about 940 basis points higher than in the second quarter of 2025. This reflects the market's recognition of the value of rapid reasoning, our continuous improvement, and the benefits of scale. Looking at the total gross margin of the core, the gross margin of core cloud and other services was 41.8%, 1,600 basis points higher than in the second quarter of 2025.

The gross margin of core hardware was 38.8%, which is 510 basis points higher than the same period last year. As we mentioned last quarter, to meet the huge demand for our fast inference services, we temporarily leased back some of our own systems from cloud customers and provided services through the Cerebras cloud. Meeting these reasoning needs as early as possible enhances our ability to meet the needs of cloud customers and grow with them over time.

We believe this will create additional long-term value for Cerebras and its shareholders. In the short term, this will reduce gross profit margin due to the higher cost of leasing back capacity. As a result, on a month-on-month basis, the core gross margin was 40.6%, compared to 46.5% in the first quarter of 2026. Had it not been for the cost increase due to leasing back more systems to increase private cloud capacity, the core gross margin would be about 500 basis points higher than now, closer to the level of the previous quarter.

Looking ahead, we expect the third quarter to be a low point of core gross margin, followed by a significant improvement in the fourth quarter of 2026 as more data centers equipped with Cerebras' own systems with lower costs are put into use. This will drive a recovery in core cloud gross margin. The core gross margin will also continue to improve in 2027, and is trending towards our target of 60% or more, for the following reasons: The market has recognized that fast tokens have higher value.

This underpins higher pricing, which will be reflected in hardware and cloud services deals that confirm revenue in the coming quarters. Over the next few quarters, we'll be phasing out the more expensive leaseback system and replacing it with a lower-cost proprietary system in our private cloud. Our product roadmap shows a 20-fold increase in throughput over the next 18 months. This reduces the cost per system and per watt of electricity production tokens. As we scale up, our supply chain bill of materials costs will be further optimized.

Since we use 5 nm nodes, our wafers cost less than those vendors that require 3 or 2 nm nodes. And we also buy a much larger amount of wafers. Finally, we don't rely on HBM, which puts pressure on manufacturers that rely on HBM to either raise prices or lose profit margins. We don't have that risk, and we are convinced that this will enhance our value proposition and pricing flexibility.

Now let's look at operating margins. Core operating losses were $33.6 million. The profit margin for core operations was negative 16%, compared to negative 42% in the same period last year, an improvement of about 2,600 basis points over the previous year. While we have more than doubled our revenue and increased our investments in various areas, we can also achieve a significant increase in our core operating profit margin, which fully demonstrates the strong operating leverage inherent in our business model.

Today, we're investing in world-class talent, manufacturing and data center capabilities, and corporate infrastructure to support the significant scale expansion we expect to achieve in the next few years. As of June 30, 2026, the remaining performance obligations were US$25.4 billion. This backlog provides us with visibility, confidence in future revenue growth, and the ability to invest ahead of time as necessary.

Our existing large strategic customers provided verification, contract visibility, and necessary financial support to build new production capacity on a large scale. At the same time, we have also been successful in expanding our accessible market and customer base. OpenAI provided us with, among other things, large-scale and cutting-edge insight. AWS provides global enterprise impact.

The recent partnership with AMD has expanded market opportunities to include analytical coupling reasoning, and fast reasoning is opening up more new markets such as the security sector. We ended the second quarter with over $8.6 billion in cash, cash equivalents, restricted cash, and marketable securities. We also have a revolving credit line of up to $850 million, which has yet to be used. Our liquidity and balance sheet conditions are strong and were further strengthened after our IPO in the second quarter.

This is a significant advantage, providing us with the flexibility and capital to invest to meet the various opportunities in these dynamic and high-growth market environments. Additionally, we also have the advantage that our net capital expenditure per megawatt is far lower than most AI cloud providers, for two key reasons: First, our capital expenditure is mainly used to deploy our own hardware in our data centers, the bill of materials costs are much lower, and it doesn't include the high profit bonuses that many other companies have to pay.

Second, our largest customer will reimburse us a significant portion of the remaining data center construction capital expenses as compensation for data center straight-through costs. Now let's take a look at our performance outlook. For the third quarter of 2026, we expect core revenue to be between $214 million and $216 million, core gross margin between 38% and 40%, and core operating margin between negative 25% and negative 23%.

For the full 2026 fiscal year, we raised our core revenue forecast to $880 million to $890 million. We raised our core gross margin forecast to 41% to 43% and raised our core operating margin forecast to minus 19% to minus 17%.

Overall, the second quarter was a very strong quarter for Cerebras for continued execution and growth. We achieved record core revenue, nearly quadrupled cloud and service revenue, and significantly exceeded our expectations in gross and operating margins. We have raised our performance guidelines for various indicators for the whole year. We've made significant progress in building capacity, capacity, and customers, and ended the quarter with over $8.6 billion in cash, cash equivalents, and investments to continue our growth plans.

We are fully prepared to more than double our revenue in 2027 and achieve significant additional growth over the next few years, while significantly increasing gross and operating margins closer to our goals. Now I'll leave the call to Andrew for the final summary.

Andrew Feldman

Co-founder, CEO, President and Chairman

Thanks, Bob. More than a decade ago, we founded Cerebras, believing we could build a better processor for AI, and that to deliver such a processor, we had to build a complete accelerator system and rack. Today, Cerebras is one of only four companies in the world — Google, Amazon, Nvidia, and Cerebras — that can independently develop processors, systems, data centers, and provide AI-based cloud services to customers. Thank you all for listening to our presentation. Well, ask the operator to open the line and enter the question and answer session.

Q&A session

operator

The first question came from UBS's Timothy Arcuri.

Timothy Arcuri

UBS Investment Banking, Research Division

Andrew, I'd like to ask a question about customer concentration. You did mention that revenue will more than triple next year. Obviously, we know that OpenAI is currently growing rapidly. So this will be a huge chunk of your current incremental revenue. I'm guessing AWS could contribute $1 billion or more next year. So what do you think of next year's customer concentration issues? For example, will these two major customers account for two-thirds of your revenue? Also, can you talk about negotiations with other companies such as Google, Microsoft, etc.?

Andrew Feldman

Co-founder, CEO, President and Chairman

Of course you can. That's a great question. I think some historical context might help to understand. In 2021, people complained that we only had government customers. Then when we won a $1 billion sovereign cloud contract, people feared we only had this one sovereign cloud customer. Then we won the biggest lab, the Frontier Lab. Then people worried we didn't have hyperscale cloud service provider customers, and then we won AWS.

So I think in every case, we can use the momentum created by the previous step to grow our business. I think OpenAI is a huge customer, and they're not only an important part of our business, but an important part of the business of every other company in this field. I think they'll still be a big part of our business next year.

But you're absolutely right, AWS and other customers, whether it's a rapidly growing code generation company or some application cases in security, etc., will account for a larger share, and OpenAI's percentage of our revenue will decline over time. But I think they'll still be an important part of our revenue next year.

Timothy Arcuri

UBS Investment Banking, Research Division

OK. Immediately after that, ask another quick follow-up question. I know, Bob, you said production capacity will increase tenfold this year over year. Can you reveal the expected increase in production capacity next year? I know Andrew said revenue will more than triple but can you talk about how much manufacturing capacity will increase year over year next year?

Robert Komin

Senior Vice President, Chief Financial Officer and Treasurer

Yes. By the end of this year, our manufacturing capacity will increase by more than ten times compared to the beginning of the period. We've signed additional facility contracts for 2027 growth, enough to increase production capacity by another three to four times, and we still have time to sign more contracts. So we're looking at huge growth over the next few years.

The next question comes from TD Cowen's Joshua Buchalter.

Joshua Buchalter

TD Cowen, Research Division

Maybe next to Tim's previous question. Can you elaborate on how the financial mechanism for dealing with Amazon will work? Will the plan be to provide the service in the AWS cloud next year, and then we'll basically see what the actual demand is, so it's hard to predict now?

Andrew Feldman

Co-founder, CEO, President and Chairman

Of course. The service will be available for use. It is deployed in Amazon's data centers. It will be available through Amazon's API service Bedrock. We are currently organizing the deployment. So I think this is roughly a framework for thinking about this issue. We expect the service to go live in the first quarter.

Joshua Buchalter

TD Cowen, Research Division

Understood. Also, regarding the partnership with AMD, can you provide more details on how to bring the joint solution to market, such as integration with Helios racks and the expected revenue generation timeline? Also, regarding this partnership with AMD, they recently acquired an inference hardware company. Can you talk about how this fits into their broader product portfolio and how it relates to your solutions?

Andrew Feldman

Co-founder, CEO, President and Chairman

Of course. There are a few points. I think we'll announce more details of the partnership arrangement with AMD in due course in the future. But I think the joint solution of putting the Helios rack before the Cerebras system, with the Helios rack being pre-filled and Cerebras responsible for decoding, is a very powerful solution — a solution where we can provide much faster speed than when using Helios alone, and at the same time provide much higher throughput than when using the Cerebras alone. This plan is very appealing, and we already have buyers.

The second question concerns AMD's recent acquisition. Look, I think the company they bought is interesting and innovative, and we love to see innovative hardware more than anyone else. I think it takes a long time to buy a hardware startup from acquisition to product delivery. We were impressed by that company's research, and I think they can find many applications in AMD's product portfolio. However, I don't think its first application scenario will be data center inference.

The next question comes from Barclays Tom O'Malley.

Kyle Bleustein

Barclays Bank, Research Division

I'm Kyle Bleustein asking questions for Tom O'Malley. I would like to continue with Josh's question about the economics of AMD transactions. I'd like to confirm that this partnership means you buy an AMD Helios rack and install it in your own cloud. Do you own all of the revenue from the customer leasing the decoupled inference solution? Or is there some sort of revenue sharing agreement?

Andrew Feldman

Co-founder, CEO, President and Chairman

Yes, the answer was as described in the first part of the question.

Kyle Bleustein

Barclays Bank, Research Division

OK. My follow-up question then is that the AWS deal is to deploy in their cloud. Have you seen a final path where Trainium and CS3 can be co-deployed in your cloud or other hyperscale cloud service provider? I'm just trying to understand how decoupled reasoning might evolve in future deployments.

Andrew Feldman

Co-founder, CEO, President and Chairman

I think we're very interested in this approach. As you can see on Google, hyperscale cloud service providers have the opportunity to deploy components they have developed outside of their own data centers. This is also a direction we are interested in, not only with AWS, but also with other cloud service providers. So I think this is very likely in future collaborations with AWS.

The next question comes from Needham & Co.'s Quinn Bolton.

Quinn Bolton

Needham & Company, LLC, Research Division

Regarding the AMD deal, I'd like to quickly confirm it. Andrew, if you buy and deploy Helios racks to the Cerebras cloud, but the service is provided to OpenAI (under your contract), does this represent an additional revenue opportunity? Or how should we look at the revenue potential in this situation? I have a follow-up question.

Andrew Feldman

Co-founder, CEO, President and Chairman

I think any time you can increase throughput while maintaining the same performance, you increase your revenue opportunities, right? Throughput is the number of customers you can support simultaneously. If you can do this without sacrificing speed, then each system can generate more revenue. I don't want to dive into the specific details of our relationship with OpenAI, though.

But one of the things that makes us so excited about this collaboration is that it preserves our speed while increasing throughput, which makes every system more profitable. Each system generates more tokens. This means that each token costs less, and each token consumes less electricity. As a result, we not only increased revenue, but also improved profit margins; not only improved profit margins and revenue, but also made every data center investment more valuable because data centers are limited by power capacity.

If you can get more tokens from a given power capacity, that can be converted into more revenue. So this is a very strong advantage. And, as AWS and AMD both partner with us to decouple, we already account for about half of the leading chip makers. So this is a very, very compelling story.

Quinn Bolton

Needham & Company, LLC, Research Division

Then, as a follow-up question, you mentioned a few times in your statement that by the end of 2027, you will increase the throughput of wafer-level engines by 20 times.

Andrew Feldman

Co-founder, CEO, President and Chairman

Correct.

Quinn Bolton

Needham & Company, LLC, Research Division

Would that eliminate the need for some decoupled computation or heterogeneous reasoning? Or would that just make the overall throughput of heterogeneous solutions faster or achieve higher throughput levels?

Andrew Feldman

Co-founder, CEO, President and Chairman

I think we're exploring ways to increase throughput, right? If your throughput increases 20 times and costs stay the same, then you're in a great position. So our systems are improving throughput. We're looking for ways to increase the throughput of our decoupling solution.

We're looking at a variety of inventions, technologies, and partnerships to continue our model of industry-leading performance and dramatically increased throughput. That's a great question. I mean, this is a central point in our roadmap thinking.

The next question comes from Morgan Stanley's Joe Moore.

Joseph Moore

Morgan Stanley, Research Division

Next to what you just mentioned, with regard to decoupling and decoding, what stage of commercialization are you currently at? We understand that you can make quick inferences at scale. You guys did it in terms of decoupling, too. Ready to deploy now? Over the next year or more, what work needs to be done to reach the level of reform you are talking about?

Andrew Feldman

Co-founder, CEO, President and Chairman

OK. I think our lab is now using GPUs to run decoupled inference. I think it will be deployed and available in Q4.

Joseph Moore

Morgan Stanley, Research Division

OK. Then you mentioned in your statement the ability to work collaboratively with deployed GPUs. So, when you work closely with Amazon and work closely with AMD, will the effects of these cooperation programs be better than what you can achieve with the Nvidia GPUs you have already deployed?

Andrew Feldman

Co-founder, CEO, President and Chairman

I think it's fair to say we haven't tried that yet. Of course, it can be said that the Helios rack will bring us a larger and better solution than using a 355 GPU. Trainium 3 will also provide a better solution than using Trainium 2. If we use other GPUs, the current top generation will undoubtedly provide better performance than the previous generation.

But I think it's also fair to say that in a decoupling solution, the performance of the previous generation GPU will also be far superior to the one without the decoupling solution, right? In an environment where everyone is trying to extend the life of their hardware and continue to generate valuable tokens, this is an important choice. Is that clear?

The next question comes from Mizuho Securities' Vijay Rakesh.

Vijay Rakesh

Mizuho Securities America Ltd., Research Division

A few quick questions. Regarding the 600 megawatts of contracted capacity and contracted electricity you mentioned, and the goal of increasing production capacity by 10 times before the end of the year, will this help accelerate the growth momentum in 2027?

Andrew Feldman

Co-founder, CEO, President and Chairman

Yes, of course. I think we're looking for data center capacity around the world every day. We're doing this because we have a huge need for quick reasoning. The faster we deploy, the faster our revenue will grow. This applies not only to our cloud business, but also to our customers' on-premise business, right? The sooner they get the data center, the faster we can ship the hardware to them. So this is our top priority, something I've invested a lot of time in.

We now have a full team, and we're pretty good at tracking data centers around the world. Also, Vijay, as you know, the work is far from over after signing the contract. We have people on site every day. We maintain communication with developers and construction companies at every stage to ensure that the project runs as planned. So I think the faster we do this, the faster we can increase our revenue.

Vijay Rakesh

Mizuho Securities America Ltd., Research Division

Understood. Then, regarding partnerships, I know you mentioned hyperscale cloud service providers, but there is also an emerging group of “new cloud service providers.” They are getting financing, and Wall Street is also developing a number of financing structures. How has this channel evolved for you?

Andrew Feldman

Co-founder, CEO, President and Chairman

OK. I think in the early days, new cloud service providers were very focused on Nvidia. I think as the business becomes more clear to investors, some new cloud service providers are diversifying and finding themselves less dependent on a single hardware vendor. So our opportunities in this category are great. Some of the new cloud service providers are multi-vendor, some use only AMD, and others are founded by people with power assets. I think in 2027 this will be an important part of our business.

operator

Thank you. No more questions at the moment. That concludes today's conference call. Thank you all for participating. You can now hang up the phone.