1809948086
Smooth Scaling: System Design for High Traffic

Advertise on podcast: Smooth Scaling: System Design for High Traffic

Rating
★★★★★
5
from
2 reviews
This podcast has
25 episodes
Language
English
Publisher
Queue-it
Explicit
No
Date created
2025/04/22
Latest episode
2026/09/01
Average duration
41 min.
Release period
23 days

Description

Smooth Scaling: System Design for High Traffic focuses on all things scalability, reliability, and performance. Tune in for expert advice on how to scale systems, control costs, boost availability, optimize performance, and get the most out of your tech stack. Host Jose Quaresma is the VP of Technical Engagement at Queue-it, working on the frontlines with some of the world’s biggest businesses on their busiest days, from Ticketmaster to Zalando to Home Office U.K. He’ll be joined by experts across industries, uncovering how major organizations design, build, and deploy systems that remain reliable at scale.

Unlock Smooth Scaling: System Design for High Traffic podcast Email contact info,
Listeners & Audience details

Email contact information

Direct podcast contact details

Listeners

Audience numbers & engagement insights

Audience details

Podcast Insights

Podcast episodes

Check latest episodes from Smooth Scaling: System Design for High Traffic podcast


How to Trust an AI Agent with AWS's Alexander Günsche and Queue-it's Moji Sarooghi
2026/09/01
When an AI agent shows up at your checkout, most systems still see a bot and block it. Alexander Guensche is a Senior Solutions Architect at AWS and the lead of TSAI, the Trust Signals for Agentic Interactions protocol that AWS is building with Trusted Shops. Moji Sarooghi is a Distinguished Product Architect at Queue-it, global leader of online traffic orchestration and managing high traffic events for the worlds biggest brands. In this episode of the Smooth Scaling Podcast, the two of them walk host Jose Quaresma through what replaces the old human-versus-bot binary. They get into why trust and identity are not the same thing, why you can have one without the other, and how a merchant can respond to an agent with something more useful than yes or no. Alex is candid about the line the protocol would not cross: identifying human users, which he describes as a slide toward totalitarian systems that rate people. The conversation closes on where this lands in practice, from reputation feedback loops to the traffic-orchestration layer Queue-it is building. A grounded look at what trust means when your visitor is software acting for someone else. Episode page --- (00:00) - Intro (01:54) - When agents started shopping, and merchants blocked them (03:13) - Millions of bots on a single product drop (05:25) - The new actor: the agent and its operator (08:15) - The privacy line: knowing without surveilling (10:33) - Automated traffic is no longer automatically bad (11:24) - The protocol landscape, and where TSAI fits (13:50) - Trust without identity: the pizza shop problem (16:58) - The model: four signal categories, five actors (23:35) - Beyond yes or no: the response spectrum (25:56) - Verifying at the edge and in the application (29:43) - Why TSAI switched to Web Bot Auth (33:47) - The line they wouldn't cross: rating humans (36:04) - Reputation, feedback loops, and gaming them (39:08) - Building it into the product, and what happens next (45:06) - Five years out: users with agents (50:00) - Rapid fire, and "trust is..." Alex is a Senior Solutions Architect at AWS with 20 years of IT experience in expert and leadership roles. He is a strong advocate of agile and DevOps practices, and he enjoys seeing serverless, cloud-native and event-driven architectures deployed at scale. He has delivered large transformation projects and successfully developed own and customers’ businesses. As an international speaker, he has held advanced technology sessions at a wide range of events. Mojtaba Sarooghi is a Distinguished Product Architect at Queue-it. Moji was one of the company’s first employees, starting his journey as a software developer over 10 years ago. He is highly experienced with AWS services, product and architectural design, managing developer teams, and defining and executing on product vision. Alexander Guensche: https://www.linkedin.com/in/alexander-guensche/Moji Sarooghi: https://www.linkedin.com/in/mojtaba-sarooghi/Host José Quaresma: https://www.linkedin.com/in/jose-quaresma/  This podcast is produced and researched by Perseu Mandillo, and brought to you by Queue-it, your virtual waiting room partner.  © Queue-it, 2026
Rebuilding Observability for the Agentic AI Era with Suman Karumuri
2026/07/07
Suman Karumuri has spent 17+ years building observability systems at Amazon, Twitter, Pinterest, Slack, and Airbnb. He was tech lead for Zipkin at Twitter, co-authored the OpenTracing specification that fed into OpenTelemetry, and has now replaced Elasticsearch twice at high-traffic platforms. In this episode of the Smooth Scaling Podcast, Suman walks host Jose Quaresma through KalDB, the open-source, cloud-native log search engine he is building into a product, which runs at petabyte scale at Slack and Airbnb. They get into why he keeps rewriting Elasticsearch instead of tuning it, how separating compute from storage on S3 changes what a log system can do, and the recovery-task trick that keeps fresh logs visible when volume spikes 10x. The back half turns to agentic AI: why agents querying logs in unpredictable bursts break traditional log stacks, why you can't sample data anymore, and what engineering leaders should measure before the bill explodes. A concrete, in-the-weeds look at running log search when the primary user is no longer human. Episode page ---(00:00) - Intro (01:04) - The origin story of KalDB (05:45) - Why keep the Elasticsearch API (07:09) - Separating compute from storage (11:07) - Solving the 12-hour log lag (14:02) - The journey of a single log message (15:56) - What makes KalDB unique (18:17) - Why not just use ClickHouse? (22:12) - The polystore: search and analytics together (23:38) - Native support for traces (24:27) - How agents change observability (31:24) - Chat as the UI, and what breaks first (35:11) - What engineering leaders should do now (39:54) - Build for a burning problem (41:15) - Rapid fire: Acquired, and "scalability is..." Suman Karumuri is the founder of KalDB, the open-source cloud-native log search engine he is now building into a product. KalDB runs at petabyte scale at Slack, Salesforce, and Airbnb. He spent 17+ years building observability systems at Amazon, Twitter, Pinterest, Slack, and Airbnb, was tech lead for Zipkin at Twitter, and co-authored the OpenTracing specification under the CNCF, which became the foundation of OpenTelemetry. He is based in San Francisco. Suman has replaced Elasticsearch twice at high-traffic platforms. And with the rise of Agentic AI querying log systems in non-deterministic bursts, this places a whole new kind of stress on those systems to perform at high scale. Suman Karumuri: https://www.linkedin.com/in/mansu/Host José Quaresma: https://www.linkedin.com/in/jose-quaresma/  This podcast is produced and researched by Perseu Mandillo, and brought to you by Queue-it, your virtual waiting room partner.  © Queue-it, 2026
Sovereign Cloud and Sovereign AI in Europe with Klaus Koefoed, CEO of T-Systems
2026/06/09
Klaus Koefoed spent more than 25 years in IT services and consultancy, at Capgemini, Deloitte, and Accenture, before taking over as CEO of T-Systems in Northern Europe. In this episode of the Smooth Scaling Podcast, Klaus walks host Jose Quaresma through the rise of sovereign cloud: why a theme almost nobody raised two years ago turned urgent in early 2025, and what actually changed. They get into Europe's real position against the American hyperscalers, why the answer is rarely either/or, and how leaders should weigh the data, operational, technology, and legal layers of sovereignty. Klaus is candid that none of this comes for free. Multi-cloud adds complexity, and the right answer depends entirely on who you are. The back half turns to sovereign AI, where T-Systems and NVIDIA have stood up a billion-euro Industrial AI Cloud, and where Europe's gigafactory ambitions, real use cases like simulated wind tunnels, and AI Scrum Teams come in. A grounded, practical look at building and running infrastructure when geopolitics is suddenly part of the architecture. Episode page ---(00:00) - Intro (00:46) - From consulting to running T-Systems (01:45) - When sovereignty went from niche to urgent (03:54) - Where Europe stands: strengths and gaps (06:50) - "Let's not be lemmings": a balanced approach (08:32) - SaaS, optionality, and freedom of movement (12:06) - The what-if scenarios CIOs miss (14:00) - The levels of sovereignty: data, ops, tech, legal (16:21) - How to actually evaluate European options (20:38) - Treat your cloud like an insurance review (23:08) - Hybrid, multi-cloud, and the move back to private (27:50) - Sovereign AI and Europe's alternatives (30:09) - Real use cases: digital twins and wind tunnels (38:07) - Gigafactories and AI Scrum Teams (40:21) - Rapid fire: resources and "scalability is..." Klaus Koefoed is CEO of T-Systems Northern Europe, leading Deutsche Telekom's B2B IT-service business across the Nordics, UK and Ireland. He joined T-Systems in June 2023 from Capgemini, where he was VP and Nordic Head of Cloud Strategy & Transformation Advisory. Before that, he was a Partner at Deloitte Consulting, with earlier years at Accenture. Copenhagen-based, with 25+ years in IT services and consultancy. His move from advising to operating arrived just as digital sovereignty stopped being a compliance footnote and became a board-level resilience question. Under Klaus, T-Systems Northern Europe has pushed hard on Sovereign Solutions for Europe and especially T Cloud Public (formerly known as Open Telekom Cloud) — now ten years old enterprise grade European Public Cloud — and recently opened the Industrial AI Cloud, Europe's largest sovereign AI-infrastructure, with 10,000 GPUs connected to the existing portfolio. That makes him one of very few executives in the region who can credibly talk about running sovereign cloud and sovereign AI at scale in Europe.  🔗 Connect Klaus Koefoed: https://www.linkedin.com/in/klauskoefoed/ Host José Quaresma: https://www.linkedin.com/in/jose-quaresma/  This podcast is researched by Joseph Thwaites, produced by Perseu Mandillo, and brought to you by Queue-it, your virtual waiting room partner.  © Queue-it, 2026
The Rise of Cloud Prem: Data Ownership in the Age of AI with Galileo's Sam Dhar
2026/05/19
Sam Dhar has spent 14 years building infrastructure at Cisco, Amazon Alexa, and Adobe, and now works as Senior Staff Engineer and AI infrastructure leader at Galileo, the enterprise AI evaluation platform. In this episode of the Smooth Scaling Podcast, Sam walks host Jose Quaresma through Cloud Prem: deploying your full product stack inside the customer's own cloud environment instead of running it as SaaS. They get into why the model is resurging, and it mostly comes down to data. Enterprises want ownership and control, plus a heavy compliance load (SOC 2, HIPAA, fully air-gapped government workloads), and they do not want a vendor sitting in the read path of their most sensitive data. Sam is candid about the hard parts. Cloud Prem can be a losing game on margins, deployment is the slowest thing in the pipeline, and every customer environment is different enough to reset the work. The conversation closes on AI: why it makes Cloud Prem urgent, the brutal GPU shortage, and why self-hosting an Opus-class model is still out of reach for most companies. A direct, practitioner-level look at where enterprise AI infrastructure is actually heading. Episode page ---(00:00) - Intro (01:08) - What Cloud Prem actually is (06:05) - Why Cloud Prem is resurging now (09:37) - Provider, vendor, customer: who owns what (11:10) - "Data is paramount": the compliance driver (14:29) - Shipping software into someone else's environment (19:57) - When Cloud Prem becomes a losing game (26:48) - Quality, and the control plane / data plane split (28:50) - Monitoring without seeing the customer's data (30:52) - Why Sam moved to AI evals (34:56) - Self-hosting LLMs and the GPU bottleneck (38:01) - Smaller runtimes, frontier-level intelligence (41:46) - Why AI makes Cloud Prem urgent (46:59) - Rapid fire: the one book to read (49:01) - "Business equals scalability" Satyam “Sam” Dhar is a senior Staff Engineer and AI infrastructure leader at Galileo, where he designs systems that support real-time LLM workflows at enterprise scale. Prior to Galileo, he spent over six years at Adobe, contributing to AI-powered product development, evaluation platforms, and large-scale data systems. Earlier in his career at Amazon, he worked on high-throughput distributed services supporting Alexa’s device orchestration. Based in San Francisco, Sam’s insights and commentary have been featured in Newsweek, CNET, InfoQ, The New Stack, The Deep View, and others. He is also a Senior Member of the Institute of Electrical and Electronics Engineers. 🔗 Connect Sam Dhar: https://www.linkedin.com/in/satyamdhar/ Host José Quaresma: https://www.linkedin.com/in/jose-quaresma/ This podcast is researched by Joseph Thwaites, produced by Perseu Mandillo, and brought to you by Queue-it, your virtual waiting room partner. © Queue-it, 2026
A Decade of Kubernetes Lessons with Chris Nesbitt-Smith
2026/04/28
Chris Nesbitt-Smith has been running Kubernetes in production since version 0.4 — long before pods, before managed services, before most of today's tooling existed. In this episode of Smooth Scaling, he sits down with José Quaresma to share what a decade of running Kubernetes for UK government citizen-facing services has taught him about scaling critical infrastructure. The conversation covers why Kubernetes was the least bad option (and largely still is), why relying on autoscaling means you've already lost, and how Gregor Hohpe's "guardrails versus lane assist" metaphor changes the way you think about capacity. Chris makes the case for climbing the service stack — SaaS first, then Functions as a Service, then Platform as a Service, and only reluctantly managed Kubernetes — and explains why tech is one of the only industries that builds critical systems without ever pricing the risk of failure. A direct, opinionated look at what scaling really demands when the stakes are real and the budget isn't infinite. Episode page ---(00:01) - Intro (01:23) - Running Kubernetes since v0.4 in UK government (04:56) - Why pod rescheduling went full circle (09:07) - "Brave and stupid": running alpha-stage K8s in production (14:58) - Helm, DevOps as a job title, and cultural drift (16:43) - Climb the service stack (SaaS → FaaS → PaaS → managed K8s) (20:48) - Why engineers resist giving up control (23:52) - Tech doesn't quantify risk the way every other industry does (27:14) - If you're relying on autoscaling, it's already too late (28:30) - The KubeCon Black Friday game: dropping requests as strategy (33:03) - Graceful degradation up the stack (35:34) - "Mostly myths": data sovereignty vs. data residency (38:35) - Cloudflare and "deploy to the world" as a different paradigm (41:53) - The legacy debt sitting in UK public sector tech (46:03) - Rapid-fire: build advice, recommended reading, scalability is... Chris Nesbitt-Smith is an independent technology strategist, a Kubernetes instructor at LearnKube, and the architect of the UK Government's National Digital Exchange. Based in London, he works at the intersection of policy, security, and modern infrastructure — advising UK and international government departments, multinational enterprises, and large NGOs on cloud-native transformation and DevSecOps. A regular speaker at KubeCon, DevSecCon, and Open Source Summit, his talks span container security, policy-as-versioned-code, and platform engineering. He also blogs regularly on his blog Cloudy with Chance of Freefall. 🔗 Connect Guest Chris Nesbitt-Smith: https://uk.linkedin.com/in/cnesbittsmith Host José Quaresma: https://www.linkedin.com/in/jose-quaresma/ This podcast is researched by Joseph Thwaites, produced by Perseu Mandillo, and brought to you by Queue-it, your virtual waiting room partner. © Queue-it, 2026
Autoscaling in Production: When It Works and When It Doesn't with Zaigham Sarfaraz and Šimon Bučko
2026/04/09
In this episode, José Quaresma sits down with two Queue-it engineers — Zaigham Sarfaraz, Engineering Manager, and Šimon Bučko, Senior Software Engineer — to talk autoscaling in production. They cover the fundamentals of horizontal and vertical scaling, why stateless architecture matters for scaling out, and what happens when the metrics you're scaling on don't match your actual bottleneck. The conversation gets real when Zaigham shares a war story of autoscaling failing during an iPhone launch — one million users in one second — and how that experience reshaped how the team thinks about pre-scaling for extreme traffic. Šimon challenges the temptation to rely on default configurations and explains why the days you most need autoscaling to work are exactly the days it might not. Episode page ---(00:00) - Introduction (00:46) - What is autoscaling under the hood? (03:25) - Why scaling down matters too (03:53) - Horizontal vs. vertical scaling (05:43) - When vertical scaling is the better choice (07:56) - Stateful vs. stateless applications (10:42) - Solving state for horizontal scaling (12:14) - The role of load balancers (14:31) - Choosing the right scaling metrics (16:46) - Is serverless the silver bullet? (21:34) - The cost paradox of autoscaling (23:40) - iPhone launch: when the whole world wants to buy a product (25:56) - Why autoscaling isn't enough for non-linear traffic (30:37) - The fallacy of the rule of thumb (32:48) - Rapid fire questions Šimon Bučko is a Senior Software Engineer at Queue-it, working across full-stack development. He is an AWS Certified Solutions Architect Professional with strong experience in software architecture and bridging the gap between business needs and technical execution.  Zaigham Sarfaraz is an Engineering Manager at Queue-it with over 15 years of experience across frontend, backend, infrastructure, and people leadership. He is an AWS Certified Cloud Practitioner and plays a key role in ensuring stable system operations while contributing to the continuous improvement of Queue-it's backend architecture.  This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2026
Observability as a Product: Building Platforms Engineers Actually Use with Iris Dyrmishi
2026/03/17
In this episode, José Quaresma speaks with Iris Dyrmishi, Senior Observability Engineer at Miro, about building an observability platform that hundreds of engineers actually trust and use. Iris explains how her team treats observability as an internal product, walks through Miro's tracing migration from Jaeger and Zipkin to OpenTelemetry with zero disruption, and shares how teams now use traces proactively to find bottlenecks before they become outages. The conversation also covers the honest downsides — alert noise, dashboard sprawl, and the cost of observability — including a recent example using eBPF and Grafana Beyla to uncover hidden networking expenses that transformed Miro's cloud bill. Episode page ---(00:00) - Intro (00:59) - Building Observability as a Product at Miro (04:08) - Migrating to OpenTelemetry (09:21) - Industry Maturity and the Business Case (12:02) - From Reactive to Proactive Observability (14:34) - Logs vs. Tracing Explained (18:04) - Team Ownership, AI, and Freedom (24:38) - The Downsides and Costs of Observability (29:58) - Rapid Fire and Close Iris Dyrmishi is a Senior Observability Engineer at Miro, where she builds and maintains the company's observability platform. She started as a backend engineer before moving into SRE roles at Worten Portugal and Farfetch, where she developed her specialty in tracing and drove OpenTelemetry migrations across large engineering organisations without disrupting existing workflows. A CNCF Ambassador, co-organiser of Kubernetes Community Days Porto, and active voice in the observability community, she writes extensively about practical adoption challenges and has spoken at KubeCon EU and on the o11ycast podcast. Her guiding philosophy: observability is a team sport. This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2026
Online Traffic in the Age of Agentic AI with Hans Skovgaard
2026/02/24
In this episode of Smooth Scaling, José Quaresma speaks with Hans Skovgaard, Chief Technology and Product Officer at Queue-it, about a shift that is already underway and accelerating fast: the internet now carries more automated bot traffic than human traffic — and agentic AI is about to make that gap much wider. Hans explains why the old model of "bots versus humans" is fundamentally broken, and why the real question is no longer who is visiting your site, but what their intent is. The conversation covers why autoscaling can no longer protect against the extreme traffic bursts that AI agents will generate, how to make bot attacks economically unviable, and what a future of AI agents buying concert tickets on your behalf actually looks like in practice. Hans also unpacks the evolving landscape of digital identity — from payment certificates to the EU Digital Identity Wallet — and what it means to build systems that can tell a genuine buyer from a scalper running 100,000 simultaneous requests. Episode page ---(00:00) - Introduction (01:19) - The Internet Just Changed — More Bots Than Humans Online (03:51) - The New Threat Isn't Bots vs. Humans. It's Intent. (06:06) - Why Autoscaling Can't Save You in the Agentic Age (09:00) - Making Attacks Expensive — The Economics of Bot Defence (11:02) - What Does the Future Actually Look Like? The AI Agent Buying Your Tickets (14:30) - The Next Generation of Challenges — Easy for Humans, Costly for Bots (18:53) - The Deeper Problem: Volatility Is Going Out of Control (20:24) - Can We Prove You're Human? Identity, Trust & the EU Wallet (25:45) - Rapid Fire (30:07) - Outro Hans J. Skovgaard is Chief Technology and Product Officer at Queue-it, the Copenhagen-founded SaaS company whose virtual waiting room technology helps the world's biggest brands manage traffic surges and prevent bot abuse during high-demand online events. With over two decades of experience leading engineering and product organisations in Nordic software companies, Hans has built a career at the intersection of deep technical expertise and strategic leadership. Before Queue-it, he served as CTPO at Penneo, a Nasdaq Copenhagen-listed RegTech company, and as CTO and VP of R&D at Capture One, where he led the company's spin-off from Phase One, launched its first SaaS product, and shipped Capture One for iPad. Earlier, he held engineering leadership roles at Milestone Systems and Microsoft. He holds an M.Sc. in Artificial Intelligence from the University of Edinburgh and an MBA from IMD, and has published research at AAAI, IEEE, and ACM.This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2026
Running High-Traffic Product Launches at Build-A-Bear with Art Huggard
2026/02/03
In this episode of Smooth Scaling, José Quaresma sits down with Art Huggard, former VP of E-Commerce at Build-A-Bear, who transformed the company's online presence from a crashing website to a $70 million business over eight years. Art shares his unconventional path from chemical engineering to e-commerce leadership at Bass Pro Shops, Hudson's Bay, and Build-A-Bear. He reveals how the company went from website crashes every hour during the 2016 holiday season to successfully managing viral product launches like Baby Yoda that sold out in four hours. Art discusses Queue-it's virtual waiting room for handling extreme traffic spikes, real-time system tuning during flash sales, and the importance of balancing technical infrastructure with guest experience. The conversation covers cloud scalability challenges, order management bottlenecks in Salesforce Commerce Cloud, and what it takes to handle 300+ orders per minute. The episode illustrates how preparation and cross-industry lessons can turn unpredictable demand into business success. Episode page ---(00:00) - Welcome to the Smooth Scaling Podcast (01:03) - From Chemical Engineer to Ecommerce Leader (05:01) - How Early Ecommerce Got the Experience Wrong (07:30) - Walking Into a Website That Was Crashing (09:56) - Why Build-A-Bear Isn't Just a Toy Company (12:03) - Using AI to Remove Bottlenecks and Ship Faster (14:43) - COVID, Baby Yoda, and Sudden Demand Spikes (16:05) - What 300 Orders a Minute Really Looks Like (22:38) - Finding the Real Bottlenecks in the Stack (25:29) - From Ammunition to Baby Yoda: Cross-Industry Lessons (27:57) - Book Recommendations and Professional Advice (30:32) - What Scalability Really Means Art Huggard is a leading expert in Digital Commerce. He has helped many well known brands such as Build-A-Bear, Bass Pro Shops, Tracker Boats, Hudson Bay and others move from chaos to High Growth. He has a keen understanding of the entire customer ecosystem including Web, Order Management, CRM, Loyalty and Digital Marketing. Known for building high performance teams Art has been an excellent mentor to many at the companies where he has worked. Most recently Art has formed Gateway-Commerce (www.gateway-commerce.com) where he provides fractional consulting to companies looking to make significant improvements to how they serve their guests.  This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo. © Queue-it, 2026
Database Scaling at Intercom: Aurora, PlanetScale & Incident Response with Engineering Director Ryan Sherlock
2026/01/13
In this episode of Smooth Scaling, José Quaresma talks with Ryan Sherlock, Director of Engineering at Intercom, about the realities of scaling databases in a fast-growing SaaS product. Ryan shares Intercom’s journey from a single MySQL database through Aurora, proxies, and per-customer scaling patterns—and what eventually pushed the team toward PlanetScale. The conversation also explores Intercom’s heartbeat-based approach to incident detection and response, focusing on customer impact rather than infrastructure metrics. Episode page ---(00:00) - Intro and episode overview (01:14) - Early scaling pains: systems going down every day (02:56) - Database evolution: MySQL, caching, Aurora, and ProxySQL (07:36) - Tens of billions of rows and the table Intercom couldn’t migrate (09:07) - Intercom’s multi-region architecture and the EU region (10:59) - Why Intercom moved from Aurora to PlanetScale (Vitess) (15:12) - PlanetScale in practice: shards, VTGate, and zero-downtime upgrades (22:39) - Heartbeat metrics and automated incident response (30:03) - AWS outage case study: DynamoDB failure and real-time recovery (34:17) - Incident mitigation lessons: “I’m now a web box” and VTGate limits (41:40) - Rapid fire questions: books, career advice, and scalability mindset Ryan Sherlock is Senior Director of Engineering at Intercom in Dublin, where he leads the core technologies and infrastructure groups that power Intercom’s AI first customer service platform. Through talks and writing on the Intercom engineering blog, he shares practical playbooks on scaling infrastructure and engineering enablement, running high leverage incident response, and using heartbeat metrics to tie reliability directly to real customer outcomes rather than just server graphs. Outside Intercom, he serves on the board of the Rails Foundation, helping steward the future of the Ruby on Rails ecosystem. Before moving into tech leadership, Ryan spent several years as a professional cyclist, an experience he wrote about in “Why you should have skin in the engineering game”, and that still shapes how he thinks about risk, ownership, and reliability in software. This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2026
Infrastructure as the Product: Designing Data-Heavy Systems with Product VP Maria Petrova
2025/12/16
Infrastructure is often treated as a backend concern, but in practice it shapes how users experience a product. In this episode of Smooth Scaling, Product VP Maria Petrova explores what it means when infrastructure becomes the product, looking at real-world, data-heavy systems where decisions around compute, data resolution, scheduling, regions, and cost directly impact scalability and user experience. The conversation dives into scaling beyond the MVP, balancing accuracy with performance, and why both engineers and product managers need to think carefully about infrastructure trade-offs when operating at scale. Episode page ---(00:00) - Welcome to the Smooth Scaling Podcast (01:02) - Infrastructure Is the Product (And Why It Shapes UX) (05:34) - Performance, Databases, and Why Compute Matters Again (07:27) - How TWAICE Scales Battery Analytics With Sensor Data (10:44) - What TWAICE Optimizes (And What It Doesn’t) (14:38) - What Product Managers Must Understand About Infrastructure (20:32) - Supermetrics: Multi-Cloud, Compliance, and Customer Expectations (25:01) - Cutting Compute Costs at TWAICE Without Losing Accuracy (32:05) - Principles for Building Scalable Data Products (34:48) - Rapid Fire: Books, Advice, and What Scalability Means Maria Petrova is a product leader known for scaling data-driven platforms and building high-performing product teams.With over a decade of experience across AdTech, eCommerce, and green tech, she’s led teams at Supermetrics, Zalando, Smartly.io, and now TWAICE, where she’s shaping AI-powered energy intelligence solutions. Maria is also the founder of Value Lab, a consultancy that embeds expert product talent into growing teams. She’s passionate about building products that truly solve customer problems at scale. This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2025
Queue-it’s Virtual Waiting Room System Design with Product Architect Moji Sarooghi
2025/12/02
In this episode, Moji Sarooghi, Distinguished Product Architect at Queue-it, breaks down the design principles and distributed systems behind Queue-it’s virtual waiting room. He explains how the team handles massive traffic spikes, upholds strict first-in, first-out fairness on request, and maintains reliability at a scale that would overwhelm most platforms. Moji also covers the shift from server-side integrations to Edge compute, how Safety Net protects against unexpected peaks, and why simplicity and failure-oriented design drive every architectural choice. A clear, technical exploration of scaling responsibly when millions depend on your system. Episode page ---(00:00) - Intro (01:34) - Visitor Flow: How the Waiting Room Works (03:24) - Edge vs. Server-Side Connectors (06:10) - Why Edge Improves Simplicity & Security (07:12) - Preventing Queue Bypass Attempts (09:14) - Connector Types & Verification Logic (12:04) - Safety Net: Automatic Peak Protection (14:54) - Scheduled Waiting Rooms + Safety Net (17:19) - FIFO at Scale (18:57) - Estimating Wait Times at Scale (20:40) - Designing for Reliability & High Traffic (24:38) - How Outflow Is Calculated (29:07) - Queue-It Token & Visitor Verification (31:02) - Cookies & Secure Access (32:35) - Key AWS Services in the Architecture (34:57) - Future: Multi-Cloud, Edge, & Bring Your Own Proxy (37:59) - Outro Mojtaba Sarooghi is a Distinguished Product Architect at Queue-it. Moji was one of the company’s first employees, starting his journey as a software developer over 10 years ago. He is highly experienced with AWS services, product and architectural design, managing developer teams, and defining and executing on product vision. This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2025
Trustpilot’s Journey From Monolith to Event-Driven with Engineering VP Angela Timofte
2025/11/18
In this episode, Angela Timofte, former VP of Global Engineering at Trustpilot, shares the decade-long journey of evolving Trustpilot’s architecture from a monolith to an event-driven, serverless-first platform. She reflects on the technical and organizational shifts that made it possible—from early trade-offs and nano-services to guardrails, templating, and chaos engineering. Angela also discusses the role of AI in engineering productivity, why staying small matters, and what scalability really means across tech, teams, and leadership. A thoughtful, candid look at modernizing systems for long-term resilience. Episode page ---(00:00) - Welcome & Episode Introduction (01:14) - What a Monolith Really Is (03:15) - Why Starting With a Monolith Made Sense (04:52) - The Breaking Point: When Scale Hit Hard (07:21) - Baby Monoliths & Early Decomposition (11:13) - The Shift to Serverless First (13:21) - Guardrails, VM Alerts & Tech Stack Choices (21:17) - Microservices to Nano Services: The Trade-offs (25:47) - Traffic Peaks, Auto-Scaling & Stress Testing (33:05) - Staying Small by Design: Team Structure & Conway’s Law (36:28) - The Impact of AI in Engineering & New Beginnings (43:28) - Rapid-Fire: Books, Advice & Defining Scalability (46:19) - Wrap-up Angela Timofte is a technology leader known for transforming organizations for scale and impact. As former VP of Global Engineering & Applied AI at Trustpilot, she led both the engineering and data science functions, driving the company’s shift from monolithic systems to scalable, event-driven, cloud-native architecture and drove a major transformation going from maintenance to value creation across the engineering organization. An AWS Serverless Hero and international speaker, she’s recognized for her work on scalability, data infrastructure, and high-performance engineering culture. Today, Angela advises companies through her consultancy, Atim Advisory, and is building a new tech venture. This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2025
Navigating ISO 27001 and Multi-Cloud with Security Architect Gabor Sivók
2025/11/04
In this episode, Gabor Sivók, Cloud Security Architect at Queue-it, shares a practical look at what it takes to secure large-scale systems in today’s cloud environments. He walks through Queue-it’s journey to ISO 27001 certification, the real trade-offs between security and performance, and how security practices adapt in multi-cloud setups. Gabor also weighs in on the growing role of AI in security operations—and why the best security work often stays invisible. A grounded conversation for anyone working at the intersection of reliability, scalability, and security. Episode page ---(00:00) - Intro & welcome to the Smooth Scaling Podcast (02:58) - AWS GameDay: Security through gamification (06:31) - The journey to ISO 27001 certification (09:34) - Balancing scalability, reliability & security (12:49) - What TLS really means for secure communication (15:29) - Moving from AWS to multi-cloud security (18:31) - How AI is changing cloud security (20:27) - The endless game of attackers vs. defenders (22:44) - Advice for starting a security career early (24:11) - Wrap-up & closing message Gabor Sivók is a Cloud Security Architect at Queue-it, where he leads security efforts across the R&D organization. With a background in infrastructure and compliance, he played a key role in Queue-it’s ISO 27001 certification and now focuses on securing multi-cloud environments at scale. Gabor works closely with platform engineering teams to embed security into architecture decisions while balancing performance, resilience, and risk. He’s also an active participant in the security community, keeping pace with emerging threats and tooling through Discord, Reddit, and bug bounty networks. This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2025
Mastering MACH Architecture & Orchestration with Sezin Cagil of Dr. Martens & the MACH Alliance
2025/10/21
In this episode, Sezin Cagil, Head of Unified Commerce Technology at Dr. Martens and MACH Alliance ambassador, shares hard-won insights on implementing and scaling MACH architecture in real-world environments. She explores how to approach modular migrations, manage complex vendor ecosystems, and prepare systems for high-traffic events. Beyond the tech, Sezin highlights the importance of team readiness, operational maturity, and aligning architecture with business needs. A must-listen for anyone navigating composable commerce or modern retail infrastructure. Episode page ---(00:00) - Intro Intro & What MACH Really Means (02:27) - How Sezin Accidentally Joined the MACH Movement (03:56) - From Monoliths to MACH: The Shift in Retail Tech (08:18) - Shopify, Complexity & When MACH Fits (11:04) - Inside a MACH E-commerce Architecture (14:52) - Aligning Tech Decisions with Business Needs (18:30) - Moving from Monolith to MACH: Benefits & Pitfalls (21:43) - Start Small: The Right Way to Transform (26:21) - Preparing for Traffic Peaks & Vendor Alignment (33:38) - The Queue-it Story: Handling Surprises in Peak Events (37:41) - Defining Scalability—Sezin’s Final Take Sezin Cagil is Head of Unified Commerce Technology at Dr. Martens, leading the teams through digital transformation and supporting omnichannel strategy. With expertise in agile delivery, composable architecture, and MACH principles, she drives seamless customer experiences across digital and retail channels. Previously, she led digital delivery at Selfridges and Costa Coffee, scaling international eCommerce platforms. As a MACH Alliance Ambassador, Sezin advocates for modern, modular technologies and contributes to industry thought leadership. This podcast is hosted by José Quaresma, researched by Joseph Thwaites and produced by Perseu Mandillo.  © Queue-it, 2025

Podcast reviews

Read Smooth Scaling: System Design for High Traffic podcast reviews


5 out of 5
2 reviews

Podcast sponsorship advertising

Start advertising on Smooth Scaling: System Design for High Traffic relevant audience podcasts


What do you want to promote?