Back to a16z Podcast

Building the Cloud for an Agentic World | AWS CEO Matt Garman...

a16z Podcast

Full Title

Building the Cloud for an Agentic World | AWS CEO Matt Garman

Summary

AWS CEO Matt Garman discusses the rapid growth of AWS, driven by AI and cloud migration, and the company's efforts to adapt its infrastructure and services for an agentic future.

Key topics include managing scarce GPU capacity, supply chain challenges, the evolution of startups, the development of agent-centric cloud services, and AWS's strategic investments in custom silicon and renewable energy.

Key Points

  • AWS continues to experience significant growth, with revenue approaching $170 billion and a 37% year-over-year increase, indicating that the cloud market is still in its early stages with substantial on-premises workloads yet to migrate.
  • Startups have evolved from small operations to highly funded, billion-dollar valuations from day one, requiring larger initial infrastructure and capital.
  • AWS is adapting its services and infrastructure to better support agents, focusing on API interfaces, low latency, and high throughput, with new services like AgentCore and AWS Context designed to facilitate agent interaction with data.
  • The company is investing heavily in custom silicon, like Graviton and Trainium, to optimize performance and cost for specific workloads, including AI training and inference.
  • Managing the demand for GPUs is a significant challenge, with AWS making intentional allocation decisions to balance large AI labs, enterprise customers, and emerging startups, despite facing supply chain constraints in power, chips, and construction.
  • AWS is rethinking traditional infrastructure paradigms for agents, considering more transient and cost-effective resource models for temporary agent tasks while maintaining durability for production systems.
  • The company is addressing the complexity of setting up new AWS accounts for agents by simplifying the onboarding process, allowing for faster deployment without immediate configuration of VPCs and IAM roles.
  • AWS's investment in its own hardware and infrastructure, including self-managed supply chains and long-term capacity planning, is crucial for meeting the accelerating demand, particularly for AI workloads.
  • The company is actively involved in developing and deploying renewable energy projects to power its data centers and is focused on being a positive force in the communities where it operates.
  • AWS is investing in agentic development internally and for its customers, aiming to transform workflows and accelerate product development by enabling agents to write code and manage tasks, while also addressing the critical need for trust and safety in autonomous agent deployments.
  • Enterprises are exploring agent adoption by rethinking workflows for parallel processing and efficiency rather than simply replicating existing processes, and are focused on building secure, autonomous agent systems with appropriate guardrails.
  • AWS is committed to protecting customer data by ensuring it never leaves their VPCs when using services like Bedrock, offering a secure environment for fine-tuning and running models, including open-source options.
  • The company is investing in AI-powered security solutions like Continuum to help customers identify and prioritize vulnerabilities at machine speed, a capability crucial for protecting complex cloud environments.
  • AWS is seeing significant adoption of its agentic capabilities across its business, with internal teams and external customers leveraging tools to accelerate software development, improve HR processes, and ensure financial compliance.

Conclusion

AWS is continuously innovating and investing to build the foundational infrastructure for an agentic world, adapting its services and hardware to meet the evolving demands of AI and cloud computing.

The company prioritizes customer data security and trust, offering robust solutions for agents and AI development while actively addressing supply chain complexities and the need for sustainable growth.

The future of cloud computing is increasingly agent-driven, with AWS poised to play a pivotal role in enabling both individuals and enterprises to harness the power of AI more effectively and safely.

Discussion Topics

  • What are the biggest challenges and opportunities in building a cloud infrastructure specifically designed for an agentic world?
  • How will the increasing prevalence of AI agents reshape the way businesses operate and the skills required for the workforce?
  • What ethical considerations and security measures are paramount as AI agents become more autonomous and integrated into critical business functions?

Key Terms

GPU
Graphics Processing Unit; specialized electronic circuits designed to rapidly manipulate and alter memory to accelerate the creation of images in a frame buffer intended for display output. Crucial for AI training and inference.
EC2
Elastic Compute Cloud; a web service that provides resizable compute capacity in the cloud.
AI
Artificial Intelligence; the simulation of human intelligence processes by machines, especially computer systems.
Agent
A piece of software that can act autonomously or semi-autonomously to perform tasks, often in response to events or instructions.
Custom Silicon
Hardware components designed and manufactured for a specific purpose or product, rather than being a general-purpose component.
Bare Metal
A physical server dedicated to a single tenant, offering maximum performance and control compared to virtualized environments.
Virtualization
The creation of a virtual version of something, such as an operating system, server, storage device, or network resources.
VPC
Virtual Private Cloud; a virtual network dedicated to a cloud account, allowing for isolation and control of network resources.
IAM
Identity and Access Management; a web service that helps you securely control access to AWS resources.
CapEx
Capital Expenditure; funds used by a company to acquire, upgrade, and maintain physical assets such as property, buildings, technology, or equipment.
HBM
High Bandwidth Memory; a type of RAM that uses a wide interface to increase memory bandwidth, often used in high-performance computing.
Inference
The process of using a trained AI model to make predictions or decisions based on new data.
Bedrock
A managed service from AWS that provides access to foundation models from various AI companies, allowing users to build and scale generative AI applications.
SageMaker
A fully managed service that provides every developer and data scientist with the ability to build, train, and deploy machine learning models quickly.
Open Weights Models
AI models where the model weights are publicly available, allowing for greater transparency, customization, and self-hosting.
Fine-tuning
The process of taking a pre-trained machine learning model and retraining it on a smaller, specific dataset to adapt it to a particular task or domain.
Post-training
Modifying a pre-trained model after its initial training to improve its performance or adapt it to new tasks, often involving data that was not part of the original training set.
Guardrails
Mechanisms or rules implemented to ensure that AI agents operate within defined boundaries and adhere to safety protocols.

Timeline

00:01:59

The discussion begins with AWS's massive revenue and growth, highlighting its early focus on startups and the ongoing migration of workloads from on-premises to the cloud.

00:03:09

The evolution of startups and their needs from AWS is discussed, noting the shift from small operations to highly funded entities.

00:06:39

The conversation transitions to how AWS is adapting its cloud services to be agent-friendly, focusing on APIs and performance for agents.

00:07:14

AWS's specific services designed for agents, like AgentCore and AWS Context, are mentioned.

00:09:04

The topic shifts to the re-architecture of AWS's core services to better support agents and simplify account setup.

00:11:49

The challenges of managing the transition to agent-centric infrastructure are explored, particularly regarding resource durability and new building blocks.

00:16:00

The critical issue of GPU availability for startups and larger companies is addressed, alongside the immense capital expenditure required.

00:17:24

Supply chain constraints, including power, chips, memory, and construction, are detailed as factors limiting capacity expansion.

00:17:42

AWS's strategy for allocating scarce GPU capacity among different customer segments is explained.

00:20:18

The acceleration of AWS's capital expenditure driven by AI demand is highlighted.

00:23:55

The conversation touches on how the scale of AWS's operations has changed its planning and processes, particularly regarding long-term infrastructure needs.

00:25:00

AWS's proactive approach to securing power and components for its data centers, including investments in renewable energy, is discussed.

00:26:53

The most severe constraints in the supply chain for 2027-2028 are considered, along with potential easing points.

00:28:32

AWS's direct management of its supply chain and its long-term strategy to anticipate and mitigate component shortages are detailed.

00:29:36

The public perception and benefits of data centers are discussed, with AWS emphasizing its positive contributions to communities.

00:32:16

The history and progress of AWS's custom silicon development, including Nitro, Graviton, and Trainium, are explored.

00:35:54

The significant benefits of AWS's Graviton processors in terms of performance and cost are highlighted.

00:36:49

The development and adoption of AWS's Trainium chips for AI training and inference are discussed, noting their high demand.

00:39:24

The adoption rate and current benefits that enterprises are seeing from agent technology are assessed.

00:40:34

Two key factors hindering broader enterprise agent adoption are identified: rethinking workflows and building trust in autonomous agents.

00:45:04

AWS's stance on customers using open-source models with their own data, particularly through Bedrock, is clarified.

00:46:13

The company's strategy for handling AI extension risk and security vulnerabilities is discussed, including the development of Continuum.

00:48:50

CEOs' concerns about AI agent safety and security are addressed, with AWS focusing on building secure deployment capabilities.

00:51:04

The current state of agent usage within AWS itself and its customer base is reviewed.

00:52:41

The biggest gains from agent adoption are identified as increased speed in software and product development.

00:53:51

Organizational changes related to agents managing people and vice versa are considered, with an emphasis on experimentation.

00:54:54

The conversation concludes with reflections on the current transformative period in technology and AWS's role in it.

Episode Details

Podcast
a16z Podcast
Episode
Building the Cloud for an Agentic World | AWS CEO Matt Garman
Published
October 8, 2026