
Agentic AI Runs Where Enterprise Software Runs: Embedded LLM Launches TokenVisor Spaces for AMD-Powered AI Clouds at AMD Advancing AI 2026
23.7.2026 20:52:21 CEST | GlobeNewswire by notified | Press release
The next AI cloud product is not a GPU hour: TokenVisor Spaces packages AMD EPYC™ CPU agent execution and AMD Instinct™ GPU inference into governed, auditable agent services
SAN FRANCISCO, July 23, 2026 (GLOBE NEWSWIRE) -- Timed withAMD Advancing AI 2026, Embedded LLM, an agentic inference infrastructure company in the AMD Instinct™ AI & HPC Software Ecosystem and a Red Hat ecosystem partner, today launched TokenVisor Spaces for AMD-powered AI clouds and enterprises.
The announcement aligns to the event's focus on AI infrastructure, enterprise deployment, developers, customers, and partners: TokenVisor Spaces is designed to help AI cloud operators turn AMD-powered systems into governed agentic AI services.
The business case is simple: agents do not only generate tokens. They need somewhere to run tools, browsers, code, files, queues, approvals, and enterprise workflows - and that work runs where enterprise software already lives: x86-64. They need policy, metering, and audit before enterprises will put them in production. And they need fast, governed access back to model inference.
TokenVisor Spaces provides exactly that on AMD server infrastructure: agent execution on AMD EPYC™ x86 CPUs, model inference on AMD Instinct GPUs, policy and metering through TokenVisor, and replayable workspaces that enterprises can audit and operate. For AI clouds, it is a way to sell more than raw GPU hours: governed agent services. For enterprises, it allows agents to work with existing applications inside isolated containers on infrastructure they control.
“The next AI cloud product is not a GPU hour. It is a governed agent service,” said Ghee Leng Ooi, CEO of Embedded LLM. “TokenVisor Spaces turns AMD-powered infrastructure into a place where agents can actually work: execute tools, persist state, follow policy, leave traces, and connect back to high-performance inference. Agentic AI runs where enterprise software runs, and that makes the CPU plus GPU platform the center of the next AI cloud.”
Partner Deployment Path
Validated on Red Hat OpenShift, TokenVisor Spaces gives enterprises isolated agent sandboxes with governed LLM access. OpenShift manages containers and GPUs; TokenVisor enforces entitlements, routing, budgets, rate limits, metering, and auditability. Embedded LLM is listed in the Red Hat Ecosystem Catalog.
“Agentic AI requires a balanced compute platform that brings CPU-based execution and GPU-accelerated inference together,” said Dan McNamara, senior vice president and general manager, Compute and Enterprise AI, AMD. “By combining AMD EPYC processors and AMD Instinct accelerators, TokenVisor Spaces shows how ecosystem software can help cloud providers and enterprises deploy open, governed agent services on infrastructure they control.”
“Modern AI infrastructure needs a balanced CPU plus GPU platform. AMD EPYC processors provide host-node performance, memory bandwidth, high-performance I/O, efficiency, and GPU ecosystem compatibility for GPU-accelerated AI systems, while AMD Instinct GPUs deliver the acceleration needed for large-scale inference. Embedded LLM's TokenVisor Spaces show how ecosystem software can bring these two planes together for production agentic AI on AMD-powered infrastructure.”
What TokenVisor Spaces Provides
TokenVisor Spaces gives each agent a controlled workspace on customer-managed infrastructure. Each capability answers something agents need that raw infrastructure does not provide:
- Agents run for days, not single requests. Spaces provides persistent x86-64 workspaces for long-running, multi-turn agents.
- Agents execute real code against real systems. Spaces provides isolated execution for shell, files, code, browser automation, and tools.
- Agents take actions enterprises cannot let run unattended. Spaces provides human approval workflows for sensitive actions.
- Agents fail in ways someone must be able to reconstruct. Spaces provides replayable event history for review, debugging, audit, and evaluation.
- Agents consume models continuously. Spaces provides governed model access through TokenVisor, including agent guardrails, policy, metering, routing, and usage controls.
Agentic inference is making data movement part of the performance path.
Long-context agents need their accumulated context back on every turn without paying full prefill each time. The market evidence is already public: SemiAnalysis reported that across 1.5M+ of its own Claude Code requests, roughly 95% of all tokens were cache reads, cutting its prompt-token bill by about 84%. Agentic inference economics are cache economics - and the cache has to live somewhere with more capacity than GPU HBM.
Embedded LLM has validated the model-plane data path for exactly this on AMD Instinct MI355X. In a production-shaped synthetic agentic replay, storage-backed KV-cache reuse using vLLM, a KV-cache management layer, AMD hipFile/GDS, and native local NVMe delivered 3.31x lower warm-turn median latency and 2.23x faster total wall-clock time versus a matched vLLM HBM prefix-cache baseline.*
Embedded LLM is collaborating with VAST Data and Tensormesh for platform-scale KV-cache reuse, agent state, trace capture, replay/evaluation, and RL data.
“KV-cache reuse is what makes long-running, multi-turn agents economically viable,” said Kuntai Du, Chief Scientist and Co-founder at Tensormesh. “By adopting LMCache, Embedded LLM brings that infrastructure to more of the AI cloud market, giving agents persistent, reusable context so they run faster, are more cost-effective, save energy and stay auditable at scale.”
“Agentic AI requires infrastructure that can efficiently manage context, data, and state across long-running AI workflows,” said Anat Heilper, Director of AI Architecture at VAST Data. “Our collaboration with Embedded LLM helps bring the VAST AI Operating System to AMD-powered AI environments, giving customers the persistent data foundation they need to scale production AI with greater performance and efficiency.”
Availability
TokenVisor Spaces is available for partner deployment and evaluation by AI cloud operators, private AI environments, and on-premises enterprises. Prospective customers can request a demo, start a trial, join an evaluation program, or arrange a proof of concept through the Embedded LLM contact form or by emailing info@embeddedllm.com. Product information is available at embeddedllm.com.
About Embedded LLM
Embedded LLM is an agentic inference infrastructure company in the AMD Instinct AI & HPC Software Ecosystem and a Red Hat ecosystem partner. The company helps AI clouds and enterprises turn GPU fleets into production AI services through vLLM-based serving, TokenVisor for governed model APIs and monetization, TokenVisor Spaces for stateful agent execution, JamAI Base for traceable AI operations, and operator-level RL/post-training infrastructure.
Media contact:
Lim Jia Qi
pr@embeddedllm.com
https://embeddedllm.com/
AMD, the AMD arrow logo, EPYC, Instinct and combinations thereof are trademarks of Advanced Micro Devices, Inc.
Technical Appendix
- Validation scope: TokenVisor Spaces deployment validation was completed on theSupermicro AS-8126GS-TNMR with AMD Instinct MI350X GPUs and dual AMD EPYC 9575F processors. The KV-cache performance results below are a separate model-plane benchmark conducted on AMD Instinct MI355X GPUs with native local NVMe.
- Benchmark scope: results measured by Embedded LLM on AMD Instinct MI355X GPUs with AMD hipFile/GDS and native local NVMe storage, using vLLM, LMCache ConnectorV1, 24 independent long-context session families, five turns per family, concurrency two, 500-token controlled outputs, Qwen3-235B-A22B-Instruct-2507-FP8, and a matched GPU-HBM prefix-cache baseline.
Detailed MI355X result:
- 3.31x lower warm-turn median latency.
- 3.33x lower warm-turn p90 latency.
- 3.32x higher warm request rate.
- 2.23x faster total wall-clock time.
- Full cache retrievals and zero fallback I/O observed.
A photo accompanying this announcement is available at https://www.globenewswire.com/NewsRoom/AttachmentNg/31b6f0c4-2977-4ab6-84ff-68eee38da222
Subscribe to releases from GlobeNewswire by notified
Subscribe to all the latest releases from GlobeNewswire by notified by registering your e-mail address below. You can unsubscribe at any time.
Latest releases from GlobeNewswire by notified
Iveco Group signs a 150 million euro term loan facility with Cassa Depositi e Prestiti to support investments in research, development and innovation11.6.2024 12:00:00 CEST | Press release
Turin, 11th June 2024. Iveco Group N.V. (EXM: IVG), a global automotive leader active in the Commercial & Specialty Vehicles, Powertrain and related Financial Services arenas, has successfully signed a term loan facility of 150 million euros with Cassa Depositi e Prestiti (CDP), for the creation of new projects in Italy dedicated to research, development and innovation. In detail, through the resources made available by CDP, Iveco Group will develop innovative technologies and architectures in the field of electric propulsion and further develop solutions for autonomous driving, digitalisation and vehicle connectivity aimed at increasing efficiency, safety, driving comfort and productivity. The financed investments, which will have a 5-year amortising profile, will be made by Iveco Group in Italy by the end of 2025. Iveco Group N.V. (EXM: IVG) is the home of unique people and brands that power your business and mission to advance a more sustainable society. The eight brands are each a
DSV, 1115 - SHARE BUYBACK IN DSV A/S11.6.2024 11:22:17 CEST | Press release
Company Announcement No. 1115 On 24 April 2024, we initiated a share buyback programme, as described in Company Announcement No. 1104. According to the programme, the company will in the period from 24 April 2024 until 23 July 2024 purchase own shares up to a maximum value of DKK 1,000 million, and no more than 1,700,000 shares, corresponding to 0.79% of the share capital at commencement of the programme. The programme has been implemented in accordance with Regulation No. 596/2014 of the European Parliament and Council of 16 April 2014 (“MAR”) (save for the rules on share buyback programmes set out in MAR article 5) and the Commission Delegated Regulation (EU) 2016/1052, also referred to as the Safe Harbour rules. Trading dayNumber of shares bought backAverage transaction priceAmount DKKAccumulated trading for days 1-25478,1001,023.01489,100,86026:3 June 20247,0001,050.597,354,13027:4 June 20245,0001,055.705,278,50028:6 June20243,0001,096.273,288,81029:7 June 20244,0001,106.174,424,68
Landsbankinn hf.: Offering of covered bonds11.6.2024 11:16:36 CEST | Press release
Landsbankinn will offer covered bonds for sale via auction held on Thursday 13 June at 15:00. An inflation-linked series, LBANK CBI 30, will be offered for sale. In connection with the auction, a covered bond exchange offering will take place, where holders of the inflation-linked series LBANK CBI 24 can sell the covered bonds in the series against covered bonds bought in the above-mentioned auction. The clean price of the bonds is predefined at 99,594. Expected settlement date is 20 June 2024. Covered bonds issued by Landsbankinn are rated A+ with stable outlook by S&P Global Ratings. Landsbankinn Capital Markets will manage the auction. For further information, please call +354 410 7330 or email verdbrefamidlun@landsbankinn.is.
Relay42 unlocks customer intelligence with a new insights and reporting module, powered by Amazon QuickSight11.6.2024 11:00:00 CEST | Press release
AMSTERDAM, June 11, 2024 (GLOBE NEWSWIRE) -- Relay42, a leading European Customer Data Platform (CDP), is leveraging Amazon QuickSight to power its new real-time customer intelligence, reporting, and dashboard module. Harnessing the breadth and quality of customer data, the new Insights module empowers marketing teams to dive deep into customer behaviors and gain invaluable insights into the performance of their marketing programs across all online, offline, paid, and owned marketing channels. Preview of the Relay42 Insights module, in pre-beta version Key capabilities of the Relay42 Insights module include: Deep insights into customer behaviors: With the Relay42 Insights module, marketers can ask unlimited questions about their data and gain a deeper understanding of how to serve their customers more effectively. Simplicity with AI-powered querying: Marketers can use artificial intelligence to query their data using natural language search, reducing the reliance on data scientists. Us
Metasphere Labs Announces X Spaces Event on the Topic of Green Bitcoin Mining and Sound Money for Sustainability11.6.2024 10:30:00 CEST | Press release
VANCOUVER, British Columbia, June 11, 2024 (GLOBE NEWSWIRE) -- Metasphere Labs Inc. (formerly Looking Glass Labs Ltd., "Metasphere Labs" or the "Company") (Cboe Canada: LABZ) (OTC: LABZF) (FRA: H1N) is thrilled to announce an engaging Twitter Spaces event on Green Bitcoin mining, energy markets, and sustainability on July 3, 2024 at 2 p.m. ET. Follow us on X at MetasphereLabs for updates and to join the event. What We'll Discuss Bitcoin Mining Basics: Understand the fundamentals of Bitcoin mining.Energy Market Dynamics: Explore how Bitcoin mining interacts with energy markets.Sustainable Innovations: Learn about our efforts to promote sustainability in Bitcoin mining.Sound Money: Discover how tamper-proof currency can enhance stability.Efficient Payment Rails: See how fast, neutral payment systems support humanitarian projects.Carbon Footprint: Compare Bitcoin's environmental impact with traditional banking. "We're excited to host this event and dive into the critical topics of Bitcoin