
AWS AI In Practice #7
About the event
We're delighted to welcome [Robert Clarke](https://www.linkedin.com/in/rjclarke7/), Principal Cloud Engineer, Cloudscaler and [Ryan Cormack](https://www.linkedin.com/in/ryancormack/), Software Engineer and Architect, Community Member
Here's what they're bringing:
Everyone's talking about serving giant LLMs - but what about serving hundreds of small models without paying to keep every one of them running? Robert worked on Cloudscaler's delivery of a production-ready AWS landing zone and inference pipeline for a large-scale clinical healthcare trial. Tonight Robert is walking us through that engagement end to end, then deep diving into Amazon SageMaker multi-model endpoints and the principles that made the platform production-ready.
MCP, ACP, AG-UI, A2A, OpenTelemetry. Agents now run on open protocols, and Ryan builds the tooling that connects them, including strands-acp, which exposes Strands agents over the Agent Client Protocol, and acp-inspector for debugging ACP traffic. Tonight Ryan is showing us how to build agents with the open-source Strands Agents SDK, run them on Amazon Bedrock AgentCore Runtime and trace every step in Amazon CloudWatch.
A big thank you to our sponsors [Cloudscaler](https://rebrand.ly/cloudscaler) & [Rayo](https://rebrand.ly/rayo-cloud) for making this event possible.
Programme: 18:00: Arrival, registration 18:15: Talks start 20:00: Networking with food and a drink provided by the generosity of our sponsors.
Session 1: *Physical AI in Surgical Logistics: Scaling Scalpel AI on Amazon SageMaker* with Robert Clarke
Modern LLMs can be terabytes in size, needing clusters of GPUs, instances and hosts to run even a single copy. But what about the other end of the scale, when you have hundreds of small ML model variations you want to serve at the same time without spending thousands running them all hot? This session explores Amazon SageMaker multi-model endpoints in the context of an engagement run by Cloudscaler to deliver a production-ready AWS landing zone and inference pipeline supporting a large-scale clinical healthcare trial. I'll walk through the engagement as a whole, then deep dive into the technical details and principles that made it so successful. You'll take away actionable best practices you can apply to your own AWS environment, and an understanding of SageMaker multi-model endpoints. Learning Takeaways
• Use Amazon SageMaker multi-model endpoints to serve hundreds of small models from a shared fleet, loading them on demand instead of running every model hot. • Judge when multi-model endpoints fit - models of similar size and framework with mixed traffic - and when cold-start latency makes a dedicated endpoint the better choice. • Apply the principles behind a production-ready AWS landing zone and inference pipeline to your own AWS environment.
About Robert Robert is a Principal Cloud Engineer at Cloudscaler, an AWS partner, where he designs and delivers secure, well-governed AWS platforms for AI and machine learning workloads. With over a decade in engineering and technical leadership across DevOps and cloud architecture, he specialises in landing zones, platform engineering and taking AI workloads from proof of concept to production.
Session 2: *Open Standards for Agents* with Ryan Cormack
In this session, we'll look at the open-source technologies available for building AI agents and the open protocols that power them. We'll focus on using AWS's open-source Strands Agents SDK to build custom agents that work with open protocols like MCP, Agent Client Protocol (ACP) and AG-UI to build interfaces for our agents. We'll look at how to run these agents on Amazon Bedrock AgentCore Runtime and monitor them with open standards like OpenTelemetry in Amazon CloudWatch. Finally, we'll look at how Agent2Agent (A2A) lets our agents talk to each other over HTTP. Learning Takeaways
• Build custom agents with the open-source Strands Agents SDK and let them talk to each other over HTTP with…




