AMD has announced the acquisition of AI chip startup Taalas, with the goal of accelerating language model inference by etching them directly into silicon. The information was released by theregister.com.

Taalas’ technology promises a significant performance leap: initial demonstrations show model-specific integrated circuits generating up to 17,000 tokens per second. This represents a radical alternative to the traditional approach of general-purpose GPUs.

Instead of running models on generic hardware, Taalas creates dedicated circuits that incorporate the model’s architecture directly into the chip. This technique, known as ‘etching models into silicon’, eliminates much of the execution overhead, resulting in lower latency and much higher throughput.

The acquisition reinforces AMD’s strategy to compete with NVIDIA in the AI accelerator market, which is currently dominated by the American giant. With Taalas, AMD gains proprietary technology that can differentiate its products in an increasingly contested landscape.

Although the financial details of the deal have not been disclosed, the move signals the growing importance of inference efficiency, especially for large-scale models that require high operational costs.

The promise of 17,000 tokens per second places Taalas far above current solutions, which typically achieve a few hundred or a few thousand tokens per second. If the technology proves out at commercial scale, it could redefine performance standards for real-time AI applications.

Experts point out that the model-specific circuit approach has limitations, such as the need to re-fabricate the chip for each new model. However, for stable and widely used models, the technique can be extremely advantageous.

AMD has not detailed the roadmap for integrating Taalas into its products, but the acquisition should accelerate the development of high-performance inference solutions. The company may also use the technology to strengthen its data center line and attract customers seeking alternatives to NVIDIA.

The news comes at a time when the race for AI efficiency is intensifying, with companies seeking to reduce costs and increase processing speed. The acquisition of Taalas is another step for AMD to position itself as a relevant player in this market.

The open-source community is also watching with interest: if Taalas’ technology is made available for open models, it could democratize access to ultra-high-performance inference, aligning with the vision of a more accessible and less concentrated AI.

For now, AMD has not commented on plans to open the technology or offer it as a service. The immediate focus seems to be integrating Taalas into the existing portfolio and exploring synergies with its GPU and CPU lines.

The acquisition also raises questions about the future of specialized chips versus general-purpose ones. While NVIDIA bets on flexible GPUs, AMD now has a card up its sleeve to compete in specific inference niches.

AMD’s move is seen as a direct response to the growing demand for low-latency inference in applications such as chatbots, virtual assistants, and autonomous systems. With Taalas, the company can offer faster and more efficient solutions for these use cases.

Expectations are that more details about the integration and the first products based on Taalas’ technology will be revealed in the coming months. Until then, the market closely follows the developments of this acquisition.

In summary, AMD’s purchase of Taalas represents a significant advance in the quest for extreme performance in AI inference, with the potential to impact the entire hardware and software ecosystem. The promise of 17,000 tokens per second is a milestone that could accelerate the adoption of language models at scale.