Ship It Weekly is a short, practical recap of what actually matters in DevOps, SRE, cloud infrastructure, and platform engineering.
Each episode, your host Brian Teller walks through the latest outages, releases, tools, and incident writeups, then translates them into “here’s what this means for your systems” instead of just reading headlines. Expect a couple of main stories with context, a quick hit of tools or releases worth bookmarking, and the occasional segment on on-call, burnout, or team culture.
This isn’t a certification prep show or a lab walkthrough. It’s aimed at people who are already working in the space and want to stay sharp without scrolling status pages, cloud updates, and blogs all week. You’ll hear about things like cloud provider incidents, Kubernetes and platform trends, Terraform and infrastructure changes, and real postmortems that are actually worth your time.
Most episodes are 15–30 minutes, so you can catch up on the way to work or between meetings. Every now and then there will be a “special” focused on a big outage or a specific theme, but the default format is simple: what happened, why it matters, and what you might want to do about it in your own environment.
If you’re the person people DM when something is broken in prod, or you’re building the cloud and platform everyone else ships on top of, Ship It Weekly is meant to be in your rotation.
Site
RSS
Apple
Données mises à jour le 02/10/2026
Classements récents
Dernières positions dans les classements Apple Podcasts et Spotify.
Liens partagés entre épisodes et podcasts
Liens présents dans les descriptions d'épisodes et autres podcasts les utilisant également.
AWS Retires DevOps Guru: What the End of Support Means, Kubernetes Cross-Namespace CVE-2026-2270, Node.js Undici WebSocket DoS & Cloudflare’s New CLI for AI Agents
Épisode 69
jeudi 1 octobre 2026 • Durée 16:41
This week on Ship It Weekly: AWS is retiring Amazon DevOps Guru and pointing customers toward CloudWatch and the newer Amazon DevOps Agent. Kubernetes disclosed a vulnerability where StatefulSet and ControllerRevision permissions can allow cross-namespace pod creation under specific conditions. A vulnerability in Undici can let a malicious WebSocket server crash a Node.js process through compressed data. And Cloudflare launched a new CLI as AI agents grow from 25 percent to 48 percent of Wrangler usage.
The bigger theme this week is how the systems around our infrastructure are changing. Managed cloud services still have lifecycles that eventually become migration work. Kubernetes authorization can depend on what controllers do with the resources users are allowed to manipulate. Applications acting as clients still process untrusted data. And infrastructure tooling is starting to treat AI agents as first-class users rather than humans who happen to automate commands.
In the lightning round: another Kubernetes vulnerability affecting Windows nodes can expose NetNTLMv2 credentials through NTLM coercion. GitHub now supports custom runners for Dependabot version and security updates. And external systems like a CMDB or internal developer portal can push repository properties into GitHub while remaining the source of truth.
And the human closer comes from Lorin Hochstein and SRE Weekly. Some availability risks are probably never going away. Resources are finite, networks fail, security controls can affect availability, and production systems have to change. Preventing individual failures still matters, but incident response is part of reliability engineering too. Sometimes improving reliability means getting better at handling the failures you cannot eliminate.
AWS Puts Elastic Beanstalk on EKS, CrowdSec Supply-Chain Breach, Critical Next.js RCE, Microsoft Disrupts EvilTokens & Why Fixing the Initial Compromise Isn’t Enough
Épisode 68
vendredi 25 septembre 2026 • Durée 17:18
This week on Ship It Weekly: AWS introduced Elastic Beanstalk Cluster Mode, allowing multiple applications to run on shared EKS infrastructure while AWS handles much of the Kubernetes complexity. CrowdSec published how a software supply-chain compromise led to attackers copying roughly 170 private repositories using a stolen OAuth token. A critical Next.js vulnerability in ImageResponse can lead to remote code execution through attacker-controlled SVG data. And Microsoft disrupted EvilTokens, a cybercrime platform linked to more than 12,000 compromised inboxes across 10,000 organizations.
The bigger theme this week is what happens after trust has been established. Elastic Beanstalk Cluster Mode puts more infrastructure behind a managed abstraction, but shared infrastructure still means understanding isolation and blast radius. CrowdSec shows how an initial compromise can become a credential problem long after the malicious code is gone. Next.js shows how something as ordinary as generating a social preview image can expose a server-side execution path. And EvilTokens shows how attackers can use valid access to move faster once inside an account.
In the lightning round: F5 has a critical BIG-IP APM vulnerability under active exploitation. GitHub Enterprise Cloud can now export an inventory of credentials with enterprise access, including PATs, SSH keys, OAuth tokens, and GitHub App credentials. Zyxel patched a vulnerability affecting GS1900 switches. And Veeam Agent for Microsoft Windows has a privilege-escalation vulnerability that can lead to SYSTEM access.
And the human closer comes back to CrowdSec. Removing the malicious package, patching the server, or reimaging the workstation does not necessarily end the incident. If an attacker already stole an OAuth token, cloud credential, SSH key, session, or registry credential, that access can survive long after the original compromise is gone. Containment means understanding not only how the attacker got in, but what they took with them
Links
AWS Elastic Beanstalk Cluster Mode
GitHub Actions Security, Cisco Email Gateway RCE, Helm 3 End-of-Life, Ubuntu 26.04 Runners & Why “Nothing Changed” Is Never the Whole Story
Épisode 67
samedi 19 septembre 2026 • Durée 15:33
This week on Ship It Weekly: GitHub Actions workflow execution protections are now generally available, giving organizations more control over who and what can trigger individual workflows. Cisco is patching critical vulnerabilities in Secure Email Gateway, including an actively exploited issue that can lead to remote command execution as root. Helm 3 has reached its final minor release and is heading toward end-of-life in February 2027. And GitHub’s ubuntu-latest Actions runner is preparing to move from Ubuntu 24.04 to 26.04.
The bigger theme this week is infrastructure that changes even when your code does not. GitHub is making CI execution permissions more explicit, Helm teams now have a defined migration deadline, and the ubuntu-latest transition is a good example of how a completely unchanged workflow can suddenly be running in a different environment. Pinning everything forever is not necessarily the answer. The important part is knowing which dependencies are allowed to move and testing those changes deliberately.
In the lightning round: GitHub Actions checks, workflow runs, and statuses will begin following your configured retention period on October 1. GitHub Advanced Security can now enforce configurations from the enterprise level. GitHub added API support for tracking when self-hosted Actions runner versions lose support. And AI Scan for pull requests can now be used without requiring CodeQL default setup.
And the human closer starts with a sentence almost every infrastructure engineer has heard during an incident: “But nothing changed.” Maybe nothing changed in the application, but the runner image changed, a dependency moved, a certificate expired, DNS changed, or an external service behaved differently. Latest tags, loose version constraints, external APIs, and even support windows are dependencies. The goal is not to freeze everything forever. It is to avoid accidental mutability, where something can change without the team realizing it was ever allowed to change.
Amazon Linux 2027, GitHub Actions Cache Security, Secret-Scanning Merge Blocks, N-central CVSS 10 RCE, Karmada Graduation, ShieldCrash, CodeQL ARM64 & When Observability Fails Too
Épisode 66
samedi 12 septembre 2026 • Durée 14:47
This week on Ship It Weekly: Amazon Linux 2027 enters public preview with kernel 7.1+, SELinux enforcing by default, DNF5, newer language runtimes, AWS-LC, and an x86-64-v3 baseline. GitHub Actions adds explicit cache permissions to reduce cache-poisoning risk. GitHub can now block pull requests from merging when they introduce exposed secrets. And N-able N-central has a critical pre-auth RCE that Huntress says is being actively exploited in the wild.
The bigger theme this week is catching problems before they turn into incidents. Amazon Linux 2027 gives teams time to test AMIs, bootstrap scripts, agents, Terraform, CloudFormation, and CI/CD before the next platform generation becomes production reality. GitHub’s new cache controls make workflow trust boundaries explicit instead of leaving them implied. And secret-scanning rulesets move credential detection directly into the merge path, where developers can actually act on it.
In the lightning round: Karmada graduates from the CNCF as multi-cluster and distributed AI scheduling grow, ShieldCrash research claims another Microsoft Defender patch bypass with SYSTEM-level access, CodeQL 2.27 adds native Linux ARM64 support, and Dependabot can now read private GitHub Packages without another personal access token.
And the human closer is about what happens when observability shares the same failure domain as the thing it is watching. A full disk is bad enough. It gets worse when logs stop writing, monitoring data disappears, and the tools used to diagnose the outage start failing too. The takeaway is not that every monitoring component needs total isolation. It is that you should know what can blind you, and make sure at least one useful signal survives the failures you care about most.
AWS GWLB TCP Reset, Azure DevOps Live Migrations to GitHub, GitHub Runner Enforcement, Docker Root Risk, Lambda IAM Updates, PostgreSQL Upgrade Traps, SonicWall Zero-Days & Better Incident Reviews
Épisode 65
vendredi 4 septembre 2026 • Durée 17:26
This week on Ship It Weekly: AWS Gateway Load Balancer gets TCP Reset, giving applications a faster way to recover when firewalls or other inline appliances fail instead of waiting minutes for TCP retries to time out. Microsoft puts Enterprise Live Migrations into public preview for moving Azure DevOps repositories to GitHub Enterprise Cloud with data residency while developers keep working. GitHub is beginning enforcement against outdated self-hosted Actions runners. And Omarchy fixes a Docker configuration that effectively gave normal desktop processes a path to root.
The bigger theme this week is failure modes hiding inside infrastructure we already trust. A dead network path can look like a slow application. A repository migration involves far more than copying Git history. A self-hosted runner can quietly become unsupported while it continues looking healthy. And giving a developer access to the Docker socket may sound like convenience until you remember that the Docker group is effectively a root-level privilege.
In the lightning round: Lambda gets full IAM resource-based policies, AWS warns that circular PostgreSQL role memberships can stall major RDS and Aurora upgrades, a researcher releases the FalconFlank CrowdStrike privilege-escalation PoC while CrowdStrike investigates, and SonicWall patches two SMA1000 zero-days after confirming active exploitation.
Cloudflare Saves 100TB of RAM, AI Drives Server Prices Up, AWS Adds a Fourth London AZ, Route 53 DNS Self-Service, AKS eBPF Routing, Go 1.27, and the Danger of Hidden Infrastructure Assumptions
Épisode 64
samedi 29 août 2026 • Durée 16:48
This week on Ship It Weekly: Cloudflare explains how five low-level optimizations to the cache behind 1.1.1.1 freed roughly 100 terabytes of RAM while also improving performance. OVHcloud is raising infrastructure prices as AI demand reshapes the memory supply chain. AWS adds a fourth Availability Zone to London, exposing automation that quietly assumed there would always be three. And Route 53 Global Resolver gets a cleaner cross-account model for DNS self-service.
The bigger theme this week is assumptions. A few wasted bytes do not matter until you have 250 billion cache entries. A Region having three Availability Zones feels permanent until AWS adds a fourth. And centralized DNS governance works fine until every application team needs a networking ticket just to make a private zone resolvable.
In the lightning round: new research looks at manipulating DRAM controller translation registers and the assumptions that creates for memory isolation, AKS eBPF Host Routing reaches general availability, CloudFront Functions can now put custom context directly into access logs, and Go 1.27 lands generic methods along with runtime, tooling, and standard-library improvements.
And the human closer looks at an easy Kubernetes mistake: running kubectl against the wrong cluster. Because the active context belongs to the kubeconfig rather than a terminal tab, changing it in one shell can silently affect another. It is a good reminder that some friction is worth keeping around production, and that the safest guardrails live somewhere stronger than operator memory.
OVHcloud Raises Prices as AI Memory Demand Reprices Non-AI Infrastructure https://tsn.io/tnaYj
AWS Adds a Fourth Availability Zone to Europe (London)
Ship It Conversations: Justin Garrison of Sidero Labs on Kubernetes, Platform Engineering, AI, Golden Paths, and Knowing What to Say No To
Épisode 63
lundi 24 août 2026 • Durée 41:10
This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.
In this Ship It Conversations episode, I talk with Justin Garrison of Sidero Labs about Kubernetes, platform engineering, bare metal, AI, golden paths, and why knowing what to say no to may be one of the most important skills a platform team can develop.
Justin is Field CTO at Sidero Labs, the company behind Talos Linux, and co-host of Fork Around and Find Out.
We start with the evolution of Kubernetes and how managed services like EKS and GKE made Kubernetes easier to consume while also pulling teams deeper into proprietary cloud ecosystems. Justin explains why on-prem and bare metal are getting renewed attention, especially as teams look at cloud costs, data sovereignty, and the operational overhead that comes with constantly optimizing cloud environments.
We also get into where Kubernetes helps and where it becomes self-inflicted pain. Justin talks about abstraction, cognitive load, and why teams tend to use familiar tools for problems they were never really designed to solve.
A big part of the conversation is platform engineering and golden paths. Justin argues that every organization needs its own path, but platforms become dangerous when they try to centralize everything. He shares why one of the best decisions his team made at Disney Plus was simply saying no to stateful workloads.
We also talk about what really belongs in a platform: security controls, logging, monitoring, software supply chain visibility, and cost management. Justin explains why centralization can help in those areas, but can become a bottleneck when applied too broadly.
Near the end, we get into AI, security, tooling dependency, and engineering culture. Justin makes the point that people have always formed strong attachments to tools, and AI is another version of that. The challenge is knowing where AI actually helps versus where it becomes another dependency teams stop questioning.
The big takeaway: good platform engineering is not about supporting everything. It is about understanding what should be standardized, what should stay flexible, and what your team should explicitly refuse to own.
GitHub Outage, PleaseFix Agentic Browser Vulnerability, AWS Certificate Manager Drops Email Validation, Cloudflare TypeScript CI Workflows, AI Observability Consolidation, and the Hidden Cost of “Simple” Platform Changes
Épisode 62
vendredi 21 août 2026 • Durée 17:39
This week on Ship It Weekly: GitHub suffers another widespread outage affecting the web interface, APIs, Actions, authentication, Copilot, and other critical developer workflows. Zenity Labs demonstrates PleaseFix attacks against agentic browsers, where malicious content can influence agents with access to authenticated sessions and privileged tools. AWS Certificate Manager is moving away from email validation, and Cloudflare is experimenting with CI pipelines defined as TypeScript instead of YAML.
The bigger theme this week is dependencies and boundaries we tend to ignore until something breaks. GitHub is no longer just where the code lives. Agentic browsers are no longer just displaying webpages. Certificate renewal is not something you want depending on someone checking an inbox. And CI pipelines have become software systems of their own.
Ship It Conversations: Ned Bellavance of Ned in the Cloud on DevOps Beyond the Buzzwords, Terraform, AI, the Future of Infrastructure as Code, and Why Fundamentals Still Matter
Épisode 61
dimanche 16 août 2026 • Durée 36:12
This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.
In this Ship It Conversations episode, I talk with Ned Bellavance of Ned in the Cloud about DevOps beyond the buzzwords, platform engineering, infrastructure as code, AI, and why fundamentals still matter even as the tools change.
Ned is the founder of Ned in the Cloud and host of the Day 2 DevOps podcast, with more than 20 years in IT across systems administration, cloud, architecture, automation, and technical education.
We start with a problem a lot of teams run into: adopting the ceremonies of DevOps without actually adopting the principles. Standups, sprints, pipelines, and tooling can make an organization look mature, but the real goal is better communication, faster feedback loops, and delivery tied to actual outcomes.
We also talk about how teams decide what to prioritize next. Security, reliability, performance, FinOps, and platform work can all matter, but chasing whatever is newest does not help if the basics are still broken.
A big part of the conversation is where infrastructure as code goes from here. We get into Terraform's state and scaling model, API rate limits, the Terraform/OpenTofu split, Terragrunt, and newer approaches like Swamp from System Initiative. AI is making infrastructure code cheaper to produce, but understanding the architecture behind that code is becoming more valuable.
That leads into learning and career development. We talk about why networking, Linux, databases, security, cloud architecture, and troubleshooting still matter, even if an LLM writes most of the syntax. Build things, get them wrong in controlled environments, troubleshoot them, and learn what is happening underneath the abstraction.
The big takeaway: tools will keep changing. Judgment, architecture, troubleshooting, and understanding the systems underneath them are much harder to automate away.
Highlights
• Why DevOps ceremony is not the same as DevOps principles
Railway US East Outage, Stripe’s Graph-Based Database Recovery, Kata Containers Host Escape, DynamoDB Vector Search, AWS Network Firewall Proxy, Gateway API 1.6, and containerd 2.4
Épisode 60
vendredi 14 août 2026 • Durée 15:52
This week on Ship It Weekly: Railway explains how an upstream network problem turned into a much larger US East outage, including storage traffic falling back onto the management network and stale connections continuing to cause problems after routing recovered. Stripe shares how graph search and state machines helped cut database pager volume by about 30 percent. Kata Containers patches a critical guest-to-host escape, and DynamoDB adds native vector search.
The bigger theme this week is what happens after the obvious failure. Fixing the route does not necessarily clear the connections created while it was broken. Automating recovery does not have to mean handing an AI agent unrestricted production access. And stronger isolation does not eliminate the components that still cross the guest-host boundary.
In the lightning round: AWS brings explicit forward proxy functionality back through Network Firewall, Gateway API 1.6 moves TCPRoute and UDPRoute to stable, and containerd 2.4 enters beta with new functionality alongside breaking changes worth finding before your next runtime upgrade.
Amazon DynamoDB now supports real-time vector search
Podcasts Similaires Basées sur le Contenu
Découvrez des podcasts liées à Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News. Explorez des podcasts avec des thèmes, sujets, et formats similaires. Ces similarités sont calculées grâce à des données tangibles, pas d'extrapolations !