Building Resilient Systems With Sam Newman
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Pragmatic Engineer podcast has published an episode with Sam Newman about designing resilient distributed systems, the trade-offs of microservices and how AI is changing software development. Newman describes resource exhaustion as a frequent cause of outages and says architecture choices, including whether to fail open or closed, depend on business context.

The Pragmatic Engineer podcast has published an episode with software architect and author Sam Newman on building resilient distributed systems, the limits of microservices and the effects of AI on software development. The discussion draws on Newman’s new book, Building Resilient Distributed Systems, and covers practical decisions such as managing outages, maintaining observability and choosing how a system should respond when a dependency fails.

Newman, author of Building Microservices, argues that microservices should not be the default choice. The episode says he calls them an architecture of “last resort.” He defines a microservice most clearly as an independently deployable service: a change to it can go live without requiring other services to be deployed at the same time. A looser definition groups services around business functions rather than technical layers.

The conversation also outlines three constraints Newman sees as central to distributed systems: information takes time to travel, a resource may be unavailable, and resources such as CPU, memory, storage and network capacity are finite. According to the episode summary, Newman says resource pools running out cause many outages he has encountered. The source does not provide incident data or a numerical estimate for that observation.

Other topics include idempotency, which prevents repeated requests from producing repeated effects, and the choice between failing open or failing closed when errors occur. The episode says that decision should reflect the business context. Newman also discusses observability and AI-related concerns, including whether specifications or code should be treated as the source of truth, and the risks he calls “cognitive debt” and “cognitive surrender.”

At a glance
announcementWhen: Published; the source provides no publi…
The developmentThe Pragmatic Engineer published a podcast episode featuring Sam Newman on resilient distributed systems, microservices and AI-assisted software development.

Resilience and Failure Choices

The episode addresses decisions that can affect whether a service remains usable during an outage and whether a transient problem affects other parts of a system. It discusses latency, unavailable dependencies and finite capacity as conditions teams may need to consider when designing safeguards, monitoring systems and planning recovery.

The discussion of microservices covers a related trade-off: independent deployment can allow teams to release changes without coordinating every service, while dividing a system into services introduces distributed-system complexity. Newman describes microservices as an architecture of “last resort” and discusses business needs and team autonomy as factors in choosing an architecture.

For readers working with AI coding tools, the episode raises questions about how teams maintain an understanding of the systems they build. The source describes Newman’s concerns and discussion of modular architecture, but does not report measured outcomes from AI adoption.

Amazon

distributed system monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Microservices to Resilience

Newman is closely associated with the history of microservices. The source recounts that, in the early 2010s, Thoughtworks colleagues James Lewis and Martin Fowler were discussing service-oriented systems whose components could be replaced quickly. At an architecture symposium in England, Lewis proposed “micro apps,” and another attendee suggested “microservices.” Lewis and Fowler published an article defining the term in March 2014; Newman published Building Microservices the following year.

The source also describes Newman’s earlier work teaching automated testing. In 2007, while working as a Thoughtworks consultant embedded at Yahoo and Google alongside people from Pivotal Labs, he and colleagues taught engineers to write automated tests and testable code. The episode’s current focus on resilience extends that practical engineering interest from testing individual software behavior to handling failures across connected services.

The Pragmatic Engineer says the episode is available on YouTube, Apple and Spotify, with a transcript and timestamps on its page. It identifies Newman’s new book as Building Resilient Distributed Systems.

“Microservices are an architecture of “last resort.””

— Sam Newman, as described in The Pragmatic Engineer episode summary

Amazon

microservices architecture books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Episode Details Still Missing

The supplied source material does not state the episode’s publication date, runtime or full guest transcript. It therefore is not possible to establish from this material when the conversation was recorded or whether the book has been released. The source also offers Newman’s observations about outage causes without incident statistics, and gives no independent evidence for the effects of AI on development practices.

The material ends during its discussion of request fingerprints for idempotency, so it does not provide the full comparison of their trade-offs. It also does not specify particular observability tools, failure-handling recommendations or case studies from the interview.

Amazon

system resilience testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Listen for the Full Discussion

Readers can listen to the episode through the YouTube, Apple and Spotify links listed by The Pragmatic Engineer and consult the transcript and timestamps on the episode page. Those materials are the next source for details beyond the supplied summary, including the complete discussion of idempotency and the examples Newman used to explain resilience and AI-related concerns.

No follow-up announcement, publication schedule or additional episode milestone is identified in the source material.

Amazon

observability tools for microservices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the episode about?

It features Sam Newman discussing resilient distributed systems, microservices, observability, idempotency and changes AI may bring to software development.

Why does Newman call microservices a last resort?

The source reports that Newman does not consider microservices a default architecture. Their defining benefit in his account is independent deployment, while adopting them also means operating a distributed system.

What three distributed-systems constraints does Newman identify?

Information takes time to travel, a resource may be unavailable, and resources such as computing power, memory, storage and network capacity are finite.

Where can listeners find the episode?

The Pragmatic Engineer lists YouTube, Apple and Spotify as listening options and says a transcript and episode timestamps are available on its page.

Source: rss

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Go Concurrency Distilled

Go Concurrency Distilled is a concise mini-book with interactive examples covering goroutines, channels, synchronization and debugging.

How “TURBINE — WERK 9” Brings SVG Motion To An AI Topic

TURBINE — WERK 9 is an interactive field guide built around a scroll-driven SVG rotor, linking page movement to the site’s industrial theme.

StreetComplete Helps Make OpenStreetMap Editing Part Of A Daily Walk

IdeaNavigator AI proposes testing a role-filtered monitor, using a StreetComplete item surfaced on Hacker News as an example.

Book Review: Is Parallel Programming Hard, And, If So, What Can You Do About It?

Search and coverage interest is spiking in Paul McKenney’s free book ‘Is Parallel Programming Hard?’ — what it covers and why the trigger is unconfirmed.