All AI News
    Discovery — Broad market AITuesday, August 25, 2026 4 min read
    AI

    Cisco and Nvidia Take AI Factories from Rack to Runtime

    Cisco & Nvidia shift AI factory focus from GPU procurement to sustained production operations and day-two lifecycle management.

    Key takeaways
    • 01Enterprises bleeding API token costs are under pressure to stand up local inference fast—but speed remains the critical failure point.
    • 02Cisco's expanded Secure AI Factory now covers full liquid-cooled rack-scale compute via Supermicro, anchored to Nvidia's Cloud Partner reference architecture to eliminate late-stage integration failures.
    • 03The partnership explicitly targets post-first-token operations: monitoring, availability, and software lifecycle.
    • 04Spectrum-X open interfaces let Cisco Silicon One switches integrate without abandoning existing tooling.
    Koko brief

    Cisco & Nvidia shift AI factory focus from GPU procurement to sustained production operations and day-two lifecycle management.

    Enterprises bleeding API token costs are under pressure to stand up local inference fast—but speed remains the critical failure point. Cisco's expanded Secure AI Factory now covers full liquid-cooled rack-scale compute via Supermicro, anchored to Nvidia's Cloud Partner reference architecture to eliminate late-stage integration failures. The partnership explicitly targets post-first-token operations: monitoring, availability, and software lifecycle. Spectrum-X open interfaces let Cisco Silicon One switches integrate without abandoning existing tooling. • **Watch:** Whether validated reference designs meaningfully compress enterprise AI factory deployment timelines below current multi-month benchmarks.

    Watch: How quickly Cisco's rack-scale validated designs reduce the gap between GPU delivery and production inference for sovereign and enterprise buyers.

    Cisco and Nvidia take AI factories from rack to runtime AI factories are moving from ambitious plans toward production, but the path from graphics processing unit acquisition to usable systems remains a race against time. Neoclouds already have customers waiting for capacity, enterprises are looking to bring inference workloads closer to home and sovereign AI programs are being built now. Those distinct buyer motions are converging around a common need: getting AI systems into production quickly enough to support the business, according to Will Eatherton (pictured, left), senior vice president of Cisco Systems Inc. “Enterprises, many of them … are spending a large amount on tokens right now,” he said. “The rush and the pressure is getting these systems up so they can start offloading what has been an [application programming interface] into using local inference. I think it’s all converging on common architectures [and] common systems, but I think speed is either what’s broken or the challenge.” Eatherton, along with Gilad Shainer (center), senior vice president of networking at Nvidia Corp., and Marc Hamilton (right), vice president of solutions architecture and engineering at Nvidia, spoke with theCUBE Research’s John Furrier at the “Cisco Secure AI Factory With Nvidia Expands to Rack Scale” event during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the rise of AI factories, rack-scale compute, Nvidia’s reference architecture, networking and the challenge of moving from the first token to sustained operations. AI factories move toward a unified deployment model Cisco is expanding its Secure AI Factory beyond its networking foundation to include full rack-scale compute. The offering brings liquid-cooled systems into a broader solution, according to Eatherton. “What we’re announcing today is that across the sovereign, enterprise and neocloud markets, we need to go big,” he said. “We are going broader with compute: We have partnered with Supermicro, bringing in the full rack scale. That is liquid-cooled, starting with Blackwell and moving to Vera Rubin. That gives us a breadth of compute systems that we can then wrap around from a Cisco standpoint, from a sales support and a software standpoint.” To make the rack-scale expansion deployable, Cisco and Nvidia are building around Cisco Validated Designs that comply with the Nvidia Cloud Partner reference architecture. The framework is intended to reduce late-stage integration problems as customers bring complex AI systems into production, according to Hamilton. “Traditional enterprises had server teams and networking teams … and the two didn’t come together until very late,” he said. “In an AI factory, because it’s a five-layer cake, everything has to work together. There are so many mistakes when customers try to go to one vendor and order networking [and] another vendor to order servers.” Getting an AI factory to its first token is only the beginning. Cisco is also focused on monitoring, availability, software upgrades and lifecycle management as customers operate these systems over time, according to Eatherton. “A lot of the industry focus is up to the point that you light up the cluster and you get your first token out,” he said. “That’s been a big focus. But the day-two aspects around monitoring, health, availability and software upgrades … are things that, from a Cisco standpoint, we’ve put a lot of focus on here over the years.” The partnership also leaves room for flexibility within the networking stack. For AI factories, that flexibility gives customers room to customize the technology without abandoning familiar operating tools. The open interfaces in Nvidia Spectrum-X allow customers to run their own technologies on top of the platform, while Spectrum-X licensing allows Cisco Silicon One switches to connect to the access network, according to Shainer. “That’s the reason that we build the reference architecture: to make sure that [customers] can get the full performance of Nvidia components,” he said. “That guarantee is now coming with the Cisco AI Factory. That’s how both of us guarantee that they’re getting the best performance out of their investment.” Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the “Cisco Secure AI Factory With Nvidia Expands to Rack Scale” event: Video 3 Image: SiliconANGLE * * A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI. Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

    Don't miss tomorrow's

    The Daily Pulse in your inbox each morning — sourced and linked.

    How often
    Keep going — across the app