Back

Explore every episode of the podcast Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News

Dive into the complete episode list for Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News. Each episode is cataloged with detailed descriptions, making it easy to find and explore specific topics. Keep track of all episodes from your favorite podcast and never miss a moment of insightful content.

Rows per page:

1–50 of 70

TitlePub. DateDuration
AWS Retires DevOps Guru: What the End of Support Means, Kubernetes Cross-Namespace CVE-2026-2270, Node.js Undici WebSocket DoS & Cloudflare’s New CLI for AI Agents01 Oct 202600:16:41

This week on Ship It Weekly: AWS is retiring Amazon DevOps Guru and pointing customers toward CloudWatch and the newer Amazon DevOps Agent. Kubernetes disclosed a vulnerability where StatefulSet and ControllerRevision permissions can allow cross-namespace pod creation under specific conditions. A vulnerability in Undici can let a malicious WebSocket server crash a Node.js process through compressed data. And Cloudflare launched a new CLI as AI agents grow from 25 percent to 48 percent of Wrangler usage.

The bigger theme this week is how the systems around our infrastructure are changing. Managed cloud services still have lifecycles that eventually become migration work. Kubernetes authorization can depend on what controllers do with the resources users are allowed to manipulate. Applications acting as clients still process untrusted data. And infrastructure tooling is starting to treat AI agents as first-class users rather than humans who happen to automate commands.

In the lightning round: another Kubernetes vulnerability affecting Windows nodes can expose NetNTLMv2 credentials through NTLM coercion. GitHub now supports custom runners for Dependabot version and security updates. And external systems like a CMDB or internal developer portal can push repository properties into GitHub while remaining the source of truth.

And the human closer comes from Lorin Hochstein and SRE Weekly. Some availability risks are probably never going away. Resources are finite, networks fail, security controls can affect availability, and production systems have to change. Preventing individual failures still matters, but incident response is part of reliability engineering too. Sometimes improving reliability means getting better at handling the failures you cannot eliminate.

Links

Amazon DevOps Guru End of Support - https://tsn.io/GQHN8

Kubernetes CVE-2026-2270: Cross-Namespace Pod Creation - https://tsn.io/BsNs8

Undici CVE-2026-85024: WebSocket Denial of Service - https://tsn.io/LncLd

Cloudflare: Introducing the cf CLI - https://tsn.io/Wk3ma

Cloudflare Forge - https://tsn.io/bAPJu

Lightning Round

Kubernetes CVE-2026-76654: Windows NTLM Coercion - https://tsn.io/36Kc3

GitHub: Custom Runners for Dependabot - https://tsn.io/tYl9K

GitHub: External Custom Properties - https://www.tellerstech.com/go/s-1fd1396d/

Human Closer

Omnipresent Availability Risks in Cloud Software - https://www.tellerstech.com/go/s-076db9dd/

Our Links

This Week’s On Call Brief - https://tsn.io/fKB9V

Ship It Weekly - https://tsn.io/NqkdP

On Call Brief - https://tsn.io/Gpz2d

AWS Puts Elastic Beanstalk on EKS, CrowdSec Supply-Chain Breach, Critical Next.js RCE, Microsoft Disrupts EvilTokens & Why Fixing the Initial Compromise Isn’t Enough25 Sep 202600:17:18

This week on Ship It Weekly: AWS introduced Elastic Beanstalk Cluster Mode, allowing multiple applications to run on shared EKS infrastructure while AWS handles much of the Kubernetes complexity. CrowdSec published how a software supply-chain compromise led to attackers copying roughly 170 private repositories using a stolen OAuth token. A critical Next.js vulnerability in ImageResponse can lead to remote code execution through attacker-controlled SVG data. And Microsoft disrupted EvilTokens, a cybercrime platform linked to more than 12,000 compromised inboxes across 10,000 organizations.

The bigger theme this week is what happens after trust has been established. Elastic Beanstalk Cluster Mode puts more infrastructure behind a managed abstraction, but shared infrastructure still means understanding isolation and blast radius. CrowdSec shows how an initial compromise can become a credential problem long after the malicious code is gone. Next.js shows how something as ordinary as generating a social preview image can expose a server-side execution path. And EvilTokens shows how attackers can use valid access to move faster once inside an account.

In the lightning round: F5 has a critical BIG-IP APM vulnerability under active exploitation. GitHub Enterprise Cloud can now export an inventory of credentials with enterprise access, including PATs, SSH keys, OAuth tokens, and GitHub App credentials. Zyxel patched a vulnerability affecting GS1900 switches. And Veeam Agent for Microsoft Windows has a privilege-escalation vulnerability that can lead to SYSTEM access.

And the human closer comes back to CrowdSec. Removing the malicious package, patching the server, or reimaging the workstation does not necessarily end the incident. If an attacker already stole an OAuth token, cloud credential, SSH key, session, or registry credential, that access can survive long after the original compromise is gone. Containment means understanding not only how the attacker got in, but what they took with them

Links

AWS Elastic Beanstalk Cluster Mode

https://tsn.io/1xaV7

CrowdSec Supply-Chain Attack Analysis

https://tsn.io/7yq2f

Next.js ImageResponse Security Advisory

https://tsn.io/8JvHp

Microsoft: Disrupting EvilTokens

https://tsn.io/DtbC9

Microsoft: EvilTokens and Device-Code Phishing

https://tsn.io/ZzwtD

F5 BIG-IP APM CVE-2026-94127

https://tsn.io/sFuKW

GitHub Enterprise Credential Inventory

https://tsn.io/7bpMn

Zyxel GS1900 Security Advisory

https://www.tellerstech.com/go/s-b2595852/

Veeam Agent for Microsoft Windows Vulnerability

https://www.tellerstech.com/go/s-166d3119/

This Week’s On Call Brief

https://tsn.io/Nnd8g

Ship It Weekly

https://tsn.io/NqkdP

On Call Brief

https://tsn.io/Gpz2d

GitHub Actions Security, Cisco Email Gateway RCE, Helm 3 End-of-Life, Ubuntu 26.04 Runners & Why “Nothing Changed” Is Never the Whole Story19 Sep 202600:15:33

This week on Ship It Weekly: GitHub Actions workflow execution protections are now generally available, giving organizations more control over who and what can trigger individual workflows. Cisco is patching critical vulnerabilities in Secure Email Gateway, including an actively exploited issue that can lead to remote command execution as root. Helm 3 has reached its final minor release and is heading toward end-of-life in February 2027. And GitHub’s ubuntu-latest Actions runner is preparing to move from Ubuntu 24.04 to 26.04.

The bigger theme this week is infrastructure that changes even when your code does not. GitHub is making CI execution permissions more explicit, Helm teams now have a defined migration deadline, and the ubuntu-latest transition is a good example of how a completely unchanged workflow can suddenly be running in a different environment. Pinning everything forever is not necessarily the answer. The important part is knowing which dependencies are allowed to move and testing those changes deliberately.

In the lightning round: GitHub Actions checks, workflow runs, and statuses will begin following your configured retention period on October 1. GitHub Advanced Security can now enforce configurations from the enterprise level. GitHub added API support for tracking when self-hosted Actions runner versions lose support. And AI Scan for pull requests can now be used without requiring CodeQL default setup.

And the human closer starts with a sentence almost every infrastructure engineer has heard during an incident: “But nothing changed.” Maybe nothing changed in the application, but the runner image changed, a dependency moved, a certificate expired, DNS changed, or an external service behaved differently. Latest tags, loose version constraints, external APIs, and even support windows are dependencies. The goal is not to freeze everything forever. It is to avoid accidental mutability, where something can change without the team realizing it was ever allowed to change.

Links

GitHub Actions Workflow Execution Protections

https://tsn.io/fbqif

Cisco Secure Email Gateway Security Advisory

https://tsn.io/jX2wk

Helm 3 End of Life

https://tsn.io/Ii7jb

Ubuntu 26.04 GitHub Actions Runners and ubuntu-latest Migration

https://tsn.io/7IJ9k

GitHub Actions Retention Changes

https://tsn.io/idFxy

GitHub Advanced Security Configuration Enforcement

https://tsn.io/8vRMx

GitHub Actions Self-Hosted Runner Lifecycle API

https://tsn.io/9UhY1

GitHub Code Scanning AI Scan

https://tsn.io/ULAVW

This Week’s On Call Brief

https://www.tellerstech.com/go/26w38/

Ship It Weekly

https://tsn.io/NqkdP

On Call Brief

https://tsn.io/Gpz2d

Amazon Linux 2027, GitHub Actions Cache Security, Secret-Scanning Merge Blocks, N-central CVSS 10 RCE, Karmada Graduation, ShieldCrash, CodeQL ARM64 & When Observability Fails Too12 Sep 202600:14:47

This week on Ship It Weekly: Amazon Linux 2027 enters public preview with kernel 7.1+, SELinux enforcing by default, DNF5, newer language runtimes, AWS-LC, and an x86-64-v3 baseline. GitHub Actions adds explicit cache permissions to reduce cache-poisoning risk. GitHub can now block pull requests from merging when they introduce exposed secrets. And N-able N-central has a critical pre-auth RCE that Huntress says is being actively exploited in the wild.

The bigger theme this week is catching problems before they turn into incidents. Amazon Linux 2027 gives teams time to test AMIs, bootstrap scripts, agents, Terraform, CloudFormation, and CI/CD before the next platform generation becomes production reality. GitHub’s new cache controls make workflow trust boundaries explicit instead of leaving them implied. And secret-scanning rulesets move credential detection directly into the merge path, where developers can actually act on it.

In the lightning round: Karmada graduates from the CNCF as multi-cluster and distributed AI scheduling grow, ShieldCrash research claims another Microsoft Defender patch bypass with SYSTEM-level access, CodeQL 2.27 adds native Linux ARM64 support, and Dependabot can now read private GitHub Packages without another personal access token.

And the human closer is about what happens when observability shares the same failure domain as the thing it is watching. A full disk is bad enough. It gets worse when logs stop writing, monitoring data disappears, and the tools used to diagnose the outage start failing too. The takeaway is not that every monitoring component needs total isolation. It is that you should know what can blind you, and make sure at least one useful signal survives the failures you care about most.

Links

Amazon Linux 2027 Public Preview

https://tsn.io/NHlEa

Amazon Linux 2027 Overview and Preview Details

https://tsn.io/izDYx

Amazon Linux 2027 Known Issues and Preview Limitations

https://tsn.io/tdugd

GitHub Actions Cache Permissions with cache-mode

https://tsn.io/8p94n

Block Pull Requests with Exposed Secrets from Merging

https://tsn.io/BspA2

N-able N-central 2026.3 Hotfix 4

https://tsn.io/xredG

Huntress: N-able N-central Vulnerability and Active Exploitation

https://tsn.io/QjNd5

Karmada Graduates from the CNCF

https://tsn.io/lcJGh

Microsoft Defender ShieldCrash Zero-Day Research

https://tsn.io/YfsJ6

CodeQL 2.27 Adds Linux ARM64 Support

https://tsn.io/lxfnF

Automatic Dependabot Access to GitHub-Hosted Registries

https://tsn.io/pSilA

Ship It Weekly

https://www.tellerstech.com/go/siw/

On Call Brief

https://www.tellerstech.com/go/ocb/

AWS GWLB TCP Reset, Azure DevOps Live Migrations to GitHub, GitHub Runner Enforcement, Docker Root Risk, Lambda IAM Updates, PostgreSQL Upgrade Traps, SonicWall Zero-Days & Better Incident Reviews04 Sep 202600:17:26

This week on Ship It Weekly: AWS Gateway Load Balancer gets TCP Reset, giving applications a faster way to recover when firewalls or other inline appliances fail instead of waiting minutes for TCP retries to time out. Microsoft puts Enterprise Live Migrations into public preview for moving Azure DevOps repositories to GitHub Enterprise Cloud with data residency while developers keep working. GitHub is beginning enforcement against outdated self-hosted Actions runners. And Omarchy fixes a Docker configuration that effectively gave normal desktop processes a path to root.

The bigger theme this week is failure modes hiding inside infrastructure we already trust. A dead network path can look like a slow application. A repository migration involves far more than copying Git history. A self-hosted runner can quietly become unsupported while it continues looking healthy. And giving a developer access to the Docker socket may sound like convenience until you remember that the Docker group is effectively a root-level privilege.

In the lightning round: Lambda gets full IAM resource-based policies, AWS warns that circular PostgreSQL role memberships can stall major RDS and Aurora upgrades, a researcher releases the FalconFlank CrowdStrike privilege-escalation PoC while CrowdStrike investigates, and SonicWall patches two SMA1000 zero-days after confirming active exploitation.

Links

AWS Gateway Load Balancer TCP Reset

https://www.tellerstech.com/go/s-d7e609ab/

Azure DevOps Enterprise Live Migrations Public Preview

https://www.tellerstech.com/go/s-ea05aff9/

GitHub Actions Self-Hosted Runner Minimum Version Enforcement

https://www.tellerstech.com/go/s-6e8540c4/

Omarchy: Any User Process Can Escalate to Root

https://www.tellerstech.com/go/s-d22971c3/

AWS Lambda Full IAM Resource-Based Policies

https://www.tellerstech.com/go/s-ff2a04b5/

Fix Circular Role Dependencies Before Upgrading RDS and Aurora PostgreSQL

https://www.tellerstech.com/go/s-e4578f52/

FalconFlank CrowdStrike Privilege Escalation PoC

https://www.tellerstech.com/go/s-8c21b00b/

SonicWall SMA1000 Zero-Day Advisory

https://www.tellerstech.com/go/s-559ffc8b/

Remote Incident Reviews: Async First, Live Later?

https://www.tellerstech.com/go/s-68ca9f5e/

This Week’s On Call Brief

https://tsn.io/L95NS

Ship It Weekly

https://www.tellerstech.com/go/siw/

On Call Brief

https://www.tellerstech.com/go/ocb/

Cloudflare Saves 100TB of RAM, AI Drives Server Prices Up, AWS Adds a Fourth London AZ, Route 53 DNS Self-Service, AKS eBPF Routing, Go 1.27, and the Danger of Hidden Infrastructure Assumptions29 Aug 202600:16:48

This week on Ship It Weekly: Cloudflare explains how five low-level optimizations to the cache behind 1.1.1.1 freed roughly 100 terabytes of RAM while also improving performance. OVHcloud is raising infrastructure prices as AI demand reshapes the memory supply chain. AWS adds a fourth Availability Zone to London, exposing automation that quietly assumed there would always be three. And Route 53 Global Resolver gets a cleaner cross-account model for DNS self-service.

The bigger theme this week is assumptions. A few wasted bytes do not matter until you have 250 billion cache entries. A Region having three Availability Zones feels permanent until AWS adds a fourth. And centralized DNS governance works fine until every application team needs a networking ticket just to make a private zone resolvable.

In the lightning round: new research looks at manipulating DRAM controller translation registers and the assumptions that creates for memory isolation, AKS eBPF Host Routing reaches general availability, CloudFront Functions can now put custom context directly into access logs, and Go 1.27 lands generic methods along with runtime, tooling, and standard-library improvements.

And the human closer looks at an easy Kubernetes mistake: running kubectl against the wrong cluster. Because the active context belongs to the kubeconfig rather than a terminal tab, changing it in one shell can silently affect another. It is a good reminder that some friction is worth keeping around production, and that the safest guardrails live somewhere stronger than operator memory.

Links

Cloudflare: How We Saved 100 Terabytes of Memory by Optimizing 1.1.1.1’s DNS Cache https://www.tellerstech.com/go/s-9d6c2943/

OVHcloud Raises Prices as AI Memory Demand Reprices Non-AI Infrastructure https://tsn.io/tnaYj

AWS Adds a Fourth Availability Zone to Europe (London) https://tsn.io/YfGxa

Shared DNS Views with Amazon Route 53 Global Resolver https://tsn.io/WEyig

DRAM Controller Register Manipulation Breaks CPU Memory Isolation https://tsn.io/QKr1x

AKS eBPF Host Routing https://tsn.io/T3PMh

CloudFront Functions Unified Logging https://tsn.io/nTXLn

Go 1.27 https://tsn.io/vfXIT

kubectl Ran on the Wrong Cluster? Fix Your Context Switching https://tsn.io/6LxfG

This Week’s On Call Brief https://tsn.io/064QE

Ship It Weekly https://www.tellerstech.com/go/siw/

On Call Brief https://www.tellerstech.com/go/ocb/

Ship It Conversations: Justin Garrison of Sidero Labs on Kubernetes, Platform Engineering, AI, Golden Paths, and Knowing What to Say No To24 Aug 202600:41:10

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It Conversations episode, I talk with Justin Garrison of Sidero Labs about Kubernetes, platform engineering, bare metal, AI, golden paths, and why knowing what to say no to may be one of the most important skills a platform team can develop.

Justin is Field CTO at Sidero Labs, the company behind Talos Linux, and co-host of Fork Around and Find Out.

We start with the evolution of Kubernetes and how managed services like EKS and GKE made Kubernetes easier to consume while also pulling teams deeper into proprietary cloud ecosystems. Justin explains why on-prem and bare metal are getting renewed attention, especially as teams look at cloud costs, data sovereignty, and the operational overhead that comes with constantly optimizing cloud environments.

We also get into where Kubernetes helps and where it becomes self-inflicted pain. Justin talks about abstraction, cognitive load, and why teams tend to use familiar tools for problems they were never really designed to solve.

A big part of the conversation is platform engineering and golden paths. Justin argues that every organization needs its own path, but platforms become dangerous when they try to centralize everything. He shares why one of the best decisions his team made at Disney Plus was simply saying no to stateful workloads.

We also talk about what really belongs in a platform: security controls, logging, monitoring, software supply chain visibility, and cost management. Justin explains why centralization can help in those areas, but can become a bottleneck when applied too broadly.

Near the end, we get into AI, security, tooling dependency, and engineering culture. Justin makes the point that people have always formed strong attachments to tools, and AI is another version of that. The challenge is knowing where AI actually helps versus where it becomes another dependency teams stop questioning.

The big takeaway: good platform engineering is not about supporting everything. It is about understanding what should be standardized, what should stay flexible, and what your team should explicitly refuse to own.

Highlights

• Why Kubernetes has become increasingly productized

• Why some teams are moving back toward on-prem and bare metal

• Where cloud cost optimization starts to become its own operational burden

• Why Kubernetes helps with abstraction and cognitive load

• Why familiar tools often get used for the wrong workloads

• What golden paths actually represent inside an organization

• Why platform teams need to know what to say no to

• What should and should not be centralized

• How AI changes engineering workflows without changing the need for judgment

• Why finding work you actually enjoy matters for avoiding burnout

Links

Sidero Labs: https://www.siderolabs.com

Talos Linux: https://www.talos.dev

Justin Garrison: https://justingarrison.com

Fork Around and Find Out: https://www.forkaroundandfindout.com

More episodes and show notes: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

LMGT Awards: https://lmgt.org

GitHub Outage, PleaseFix Agentic Browser Vulnerability, AWS Certificate Manager Drops Email Validation, Cloudflare TypeScript CI Workflows, AI Observability Consolidation, and the Hidden Cost of “Simple” Platform Changes21 Aug 202600:17:39

This week on Ship It Weekly: GitHub suffers another widespread outage affecting the web interface, APIs, Actions, authentication, Copilot, and other critical developer workflows. Zenity Labs demonstrates PleaseFix attacks against agentic browsers, where malicious content can influence agents with access to authenticated sessions and privileged tools. AWS Certificate Manager is moving away from email validation, and Cloudflare is experimenting with CI pipelines defined as TypeScript instead of YAML.

The bigger theme this week is dependencies and boundaries we tend to ignore until something breaks. GitHub is no longer just where the code lives. Agentic browsers are no longer just displaying webpages. Certificate renewal is not something you want depending on someone checking an inbox. And CI pipelines have become software systems of their own.

Links

GitHub Hit by Widespread Outage https://devops.com/github-hit-by-widespread-outage-halting-work-for-global-developers/

Zenity Labs: PleaseFix in Agentic Browsers https://zenity.io/company-overview/newsroom/company-news/zenity-labs-exposes-the-full-scope-of-pleasefix

AWS Certificate Manager Ending Email Validation https://aws.amazon.com/blogs/security/aws-certificate-manager-will-discontinue-email-validation-to-prove-domain-validation-for-certificates/

Certificate Expiry Is Still Taking Down Major Platforms https://tokentimer.ch/blog/tls-certificate-expiry-outages

Cloudflare Turns CI Pipelines into TypeScript Workflows https://www.infoq.com/news/2026/08/cloudflare-ci-code-workflows/

Dynatrace Acquires Arize https://devops.com/dynatrace-acquires-arize-as-ai-agents-deepen-the-observability-challenge/

AWS Open-Sources Dogwood https://www.infoq.com/news/2026/08/aws-dogwood-agent-policy/

Pulumi v3.258.0 https://github.com/pulumi/pulumi/releases/tag/v3.258.0

AWS Key Breach and Data-Transfer Signal https://assets.theregister.com/2026/08/13/20267/

Mario Saved the EU but Broke My System https://www.uptimelabs.io/articles/hamed-2012-outage-reflections

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W34/

Ship It Weekly https://shipitweekly.fm/

Ship It Conversations: Ned Bellavance of Ned in the Cloud on DevOps Beyond the Buzzwords, Terraform, AI, the Future of Infrastructure as Code, and Why Fundamentals Still Matter16 Aug 202600:36:12

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It Conversations episode, I talk with Ned Bellavance of Ned in the Cloud about DevOps beyond the buzzwords, platform engineering, infrastructure as code, AI, and why fundamentals still matter even as the tools change.

Ned is the founder of Ned in the Cloud and host of the Day 2 DevOps podcast, with more than 20 years in IT across systems administration, cloud, architecture, automation, and technical education.

We start with a problem a lot of teams run into: adopting the ceremonies of DevOps without actually adopting the principles. Standups, sprints, pipelines, and tooling can make an organization look mature, but the real goal is better communication, faster feedback loops, and delivery tied to actual outcomes.

We also talk about how teams decide what to prioritize next. Security, reliability, performance, FinOps, and platform work can all matter, but chasing whatever is newest does not help if the basics are still broken.

A big part of the conversation is where infrastructure as code goes from here. We get into Terraform's state and scaling model, API rate limits, the Terraform/OpenTofu split, Terragrunt, and newer approaches like Swamp from System Initiative. AI is making infrastructure code cheaper to produce, but understanding the architecture behind that code is becoming more valuable.

That leads into learning and career development. We talk about why networking, Linux, databases, security, cloud architecture, and troubleshooting still matter, even if an LLM writes most of the syntax. Build things, get them wrong in controlled environments, troubleshoot them, and learn what is happening underneath the abstraction.

The big takeaway: tools will keep changing. Judgment, architecture, troubleshooting, and understanding the systems underneath them are much harder to automate away.

Highlights

• Why DevOps ceremony is not the same as DevOps principles

• How teams should decide what platform work actually matters

• Why basic security hygiene still matters

• Where Terraform's current state model starts to hit limits

• Terraform, OpenTofu, Terragrunt, and the future of infrastructure as code

• How AI changes the value of writing infrastructure code

• Why architecture and troubleshooting skills become more important

• Why breaking things in controlled environments is one of the best ways to learn

Links

Ned in the Cloud: https://nedinthecloud.com

Day 2 DevOps: https://day2devops.com

Ned in the Cloud on YouTube: https://www.youtube.com/c/NedintheCloud/

Ned Bellavance on LinkedIn: https://www.linkedin.com/in/ned-bellavance/

Swamp: https://swamp.club

Terraform: https://developer.hashicorp.com/terraform

OpenTofu: https://opentofu.org

More episodes and show notes: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

Railway US East Outage, Stripe’s Graph-Based Database Recovery, Kata Containers Host Escape, DynamoDB Vector Search, AWS Network Firewall Proxy, Gateway API 1.6, and containerd 2.414 Aug 202600:15:52

This week on Ship It Weekly: Railway explains how an upstream network problem turned into a much larger US East outage, including storage traffic falling back onto the management network and stale connections continuing to cause problems after routing recovered. Stripe shares how graph search and state machines helped cut database pager volume by about 30 percent. Kata Containers patches a critical guest-to-host escape, and DynamoDB adds native vector search.

The bigger theme this week is what happens after the obvious failure. Fixing the route does not necessarily clear the connections created while it was broken. Automating recovery does not have to mean handing an AI agent unrestricted production access. And stronger isolation does not eliminate the components that still cross the guest-host boundary.

In the lightning round: AWS brings explicit forward proxy functionality back through Network Firewall, Gateway API 1.6 moves TCPRoute and UDPRoute to stable, and containerd 2.4 enters beta with new functionality alongside breaking changes worth finding before your next runtime upgrade.

Links

Railway: July 2, 2026 US East Services Outage https://blog.railway.com/p/incident-report-july-2-2026-us-east-services-outage

Stripe: How Stripe uses graph search and state machines to auto-remediate a global database fleet https://stripe.dev/blog/how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-global-database-fleet

Kata Containers: Guest-root to host-root escape via virtiofs https://github.com/kata-containers/kata-containers/security/advisories/GHSA-2gv2-cffp-j227

Amazon DynamoDB now supports real-time vector search https://aws.amazon.com/about-aws/whats-new/2026/08/amazon-dynamodb-vector-search/

AWS Network Firewall forward proxy preview https://aws.amazon.com/about-aws/whats-new/2026/08/aws-network-firewall-forward-proxy-preview/

Gateway API v1.6: TCPRoute and UDPRoute Graduate to Standard https://kubernetes.io/blog/2026/08/03/gateway-api-v1-6-release/

containerd 2.4 beta https://github.com/containerd/containerd/releases

CNCF: Learning Cloud-Native Engineering Beyond Tutorials Through LFX https://www.cncf.io/blog/2026/08/10/learning-cloud-native-engineering-beyond-tutorials-through-lfx/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W33/

Ship It Weekly https://shipitweekly.fm/

AI Agents Target GitHub, Kubernetes 1.37 Deprecations, AWS Transit Gateway Policy Routing, IAM Identity Center Multi-Region, Cloudflare Meerkat, AMOS Mac Malware, and the Risk of Half-Migrated Production Systems07 Aug 202600:18:19

This week on Ship It Weekly: AI agents from Anthropic and OpenAI took unsanctioned actions on the real internet during UK government cyber testing, including an attempt to push malicious code into a real GitHub project. Kubernetes 1.37 starts retiring IPVS mode, pushes cgroup v1 closer to removal, and brings an SELinux volume change worth testing before upgrades. AWS Transit Gateway gets policy-based routing, and IAM Identity Center expands multi-Region support to organizations using AWS’s built-in directory.

The bigger theme: access is not the same thing as authority, and availability is not just about whether your application is running. Agents need boundaries around the actions they can take. Routing policies need enough visibility to explain why traffic went where it did. And regional resilience does not help much if the people responding to the outage cannot authenticate.

In the lightning round: Cloudflare’s Meerkat consensus system, AMOS macOS malware, N-able’s incomplete N-central fix, and an AWS CLI bug that disabled SSH host-key verification.

Links

AISI: Unsanctioned agent behaviour during cyber testing https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing

Kubernetes v1.37 Sneak Peek https://kubernetes.io/blog/2026/07/31/kubernetes-v1-37-sneak-peek/

AWS Transit Gateway Policy-Based Routing https://aws.amazon.com/about-aws/whats-new/2026/07/aws-transit-gateway-policy-based-routing/

AWS IAM Identity Center multi-Region directory support https://aws.amazon.com/about-aws/whats-new/2026/07/aws-iam-identity-center-extends-multi-region-support-to-identity-center-directory

Cloudflare: Introducing Meerkat https://blog.cloudflare.com/meerkat-introduction/

Atomic macOS / AMOS stealer infection https://isc.sans.edu/diary/rss/33208

N-able N-central exploitation after incomplete fix https://thehackernews.com/2026/08/n-able-says-attackers-take-over-n.html

CVE-2026-18654: AWS CLI EMR SSH host-key verification https://aws.amazon.com/security/security-bulletins/rss/2026-071-aws/

CloudFront VPC Origins half-migrated incident https://www.reddit.com/r/devops/comments/1vdovj6/the_cloudfront_vpc_origins_outage_caught_me/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W32/

More episodes https://shipitweekly.fm/

Telstra’s Time Sync Outage, DoorDash’s 1.5M RPS Cache, Stateless MCP, GitHub and PyPI Supply Chain Delays, and Why the Quietest Infrastructure Often Has the Biggest Blast Radius31 Jul 202600:16:57

This week on Ship It Weekly: Telstra’s mobile network jumped back to 2006 after a timing device restarted with the wrong date, disrupting calls, data sessions, and hundreds of emergency calls.

DoorDash explains how Entity Cache, built with Envoy and Valkey, handles more than 1.5 million requests per second and uses stale-data policies, invalidation, and fallback behavior as a reliability layer.

The latest MCP release candidate removes protocol-level sessions, making servers easier to scale behind ordinary load balancers while leaving teams responsible for authentication, tracing, retries, rate limits, and application state.

GitHub and PyPI are also adding friction to package automation. Dependabot now delays routine updates by three days, while PyPI blocks new files from releases older than 14 days.

Links

Telstra outage https://www.telstra.com.au/exchange/our-mobile-network-outage-has-been-resolved-heres-what-happened

DoorDash Entity Cache https://careersatdoordash.com/blog/high-performance-proxy-cache-for-doordash-services/

MCP specification release candidate https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/

Dependabot package cooldown https://github.blog/changelog/2026-07-14-dependabot-version-updates-introduce-default-package-cooldown/

PyPI release-file restrictions https://blog.pypi.org/posts/2026-07-22-releases-now-reject-new-files-after-14-days/

Amazon ECS Action Logs https://aws.amazon.com/about-aws/whats-new/2026/07/amazon-ecs-action-logs/

Network Load Balancer listener rules https://aws.amazon.com/about-aws/whats-new/2026/07/aws-network-load-balancer-supports-listener-rules/

Amazon Managed Prometheus limits https://aws.amazon.com/about-aws/whats-new/2026/07/amazon-managed-service-prometheus-1500m-metrics-workspace/

PixelSmash in FFmpeg https://jfrog.com/blog/pixelsmash-critical-ffmpeg-vulnerability-turns-media-files-into-weapons/

SRE Weekly Issue 527 https://sreweekly.com/sre-weekly-issue-527/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W31/

More episodes https://shipitweekly.fm/

Ship It Conversations: Jay Lark of Hookbridge on Webhook Reliability, Retries, Idempotency, Replay, HMAC Security, Local Testing, and Operating Webhooks in Production27 Jul 202600:31:26

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It Conversations episode, I talk with Jay Lark of Hookbridge about webhook reliability, retries, idempotency, replay, security, local development, and what happens when a simple HTTP POST becomes production infrastructure.

Jay is a Principal DevOps Engineer and the founder of Hookbridge, a service focused on making webhook delivery more reliable and easier to operate.

We start with the basic model: one service sends an HTTP POST to another when something happens. The happy path is easy. The problems begin when an endpoint is unavailable, responds too slowly, receives duplicate events, or gets them out of order.

Jay explains what “at least once delivery” means and why receivers must expect duplicates. We talk about idempotency, event IDs, retry behavior, availability during deployments, and why returning a 200 does not prove downstream processing succeeded.

We also dig into observability and security. Teams need enough visibility to know whether a webhook arrived, whether signature verification passed, what response was returned, and where processing failed. Jay breaks down HMAC signatures, timestamps, replay protection, and why a valid signature still does not replace normal business-logic validation.

Local development is another source of friction. External providers cannot send events directly to localhost, so developers often rely on temporary tunnels, staging deployments, copied payloads, or mocks. Jay explains how Hookbridge uses a fixed URL and local client to forward real webhook traffic to a developer’s machine.

We also talk about n8n, self-hosted OpenClaw systems, Hookbridge pull endpoints, and when polling may be simpler or safer than exposing another inbound endpoint.

The big takeaway: design the failure path before a webhook becomes business-critical. Verify the sender, expect duplicate and out-of-order events, build enough visibility to debug failures, and have a replay strategy before the first incident.

Highlights

• Why webhooks are harder than “just an HTTP POST”

• What at-least-once delivery means for receivers

• Why idempotency, retries, and event ordering matter

• What teams need for webhook observability and debugging

• Where HMAC signatures, timestamps, and replay protection fit

• Why local webhook development is still awkward

• When polling or pull-based delivery may be a better fit

• When teams should stop building webhook infrastructure themselves

Links

Hookbridge: https://hookbridge.io

Hookbridge local development CLI: https://www.hookbridge.io/cli.html

Hookbridge pull endpoints: https://www.hookbridge.io/pull.html

Jay Lark on LinkedIn: https://www.linkedin.com/in/jay-lark-ba7a3b5/

OpenClaw: https://openclaw.ai

More episodes and show notes: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

AWS CloudFormation Express Mode, Spark 4.2 Vector Search, AI Speeds Coding but Not Delivery, Platform Governance Developers Won’t Hate, and Why Could Is Not the Same as Should24 Jul 202600:18:34

This week on Ship It Weekly: AWS CloudFormation Express mode promises faster infrastructure feedback by reporting deployments complete before extended resource stabilization finishes. Apache Spark 4.2 adds native vector operations and nearest-neighbor joins, giving some teams a way to keep AI data workloads closer to the platforms they already run.

GitLab’s latest research says AI is helping developers generate and commit code faster, but review, testing, governance, and deployment are not accelerating at the same pace. Then, a platform engineering case study from Sevdesk shows how minimum viable governance, useful feedback, and progressive enforcement can improve compliance without turning the platform team into another approval queue.

The theme this week: making one part of the system faster does not automatically improve the whole system.

In the lightning round: OpenShift 4.22.5 receives an important security update, outdated autoscaling thresholds disrupt GitHub Actions, Cloudflare experiences an incident where some POST requests fail to reach customer origins, and GitHub Code Quality becomes generally available—with automatic billing attached.

The episode closes with Reid Savage reflecting on their first year managing Honeycomb’s SRE team and the difference between what a capable team could do and what it should do.

Links

CloudFormation Express mode https://aws.amazon.com/blogs/aws/accelerate-your-infrastructure-deployments-by-up-to-4x-with-aws-cloudformation-express-mode/

Apache Spark 4.2 https://spark.apache.org/releases/spark-release-4-2-0.html

GitLab AI Accountability Report https://about.gitlab.com/resources/ai-accountability-survey-2026/

Platform governance at Sevdesk https://www.infoq.com/presentations/platform-engineering-team-compliance/

OpenShift 4.22.5 security update https://access.redhat.com/errata/RHSA-2026%3A37585

GitHub Actions incident https://github.com/orgs/community/discussions/201795

Cloudflare status https://www.cloudflarestatus.com/

GitHub Code Quality GA https://github.blog/changelog/2026-07-20-github-code-quality-is-now-generally-available/

Could vs. Should — Reid Savage https://www.honeycomb.io/blog/could-should-first-year-managing-sre-team

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W30/

More episodes and full show notes https://shipitweekly.fm/

Ship It Conversations: Mat Ryer of Grafana Labs on AI Observability, Agents, Evals, and Operating AI in Production20 Jul 202600:45:06

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It Conversations episode, I talk with Mat Ryer of Grafana Labs about AI observability, production agents, evals, telemetry cost, guardrails, and what changes once AI moves beyond demos and into systems teams actually depend on.

Mat is Senior Director of AI at Grafana Labs, where he focuses on how AI fits into observability and production systems.

We talk about Grafana Assistant, why AI observability is not just logs, latency, and HTTP 200s, and how teams can measure whether agents are actually helping. Mat gets into evals, LLM-as-judge patterns, traces as a way to think about conversations, user feedback, tool choice, model changes, and the cost of collecting new telemetry.

We also dig into UX and trust. If an AI assistant gives you a wall of text, you still have to decide whether to believe it. If it can show the graph, deep link into Grafana, apply filters, and expose the source data, that becomes a much more useful operating experience.

The big takeaway: start small, enhance workflows you already have, build feedback loops, and treat production AI like something you actually have to operate.

Highlights

• Why AI demos are easy, but production AI is harder

• Why agents need observability, evals, and guardrails

• How Grafana thinks about AI Assistant and AI observability

• Where LLM-as-judge patterns, traces, tool calls, and feedback fit

• Why telemetry cost problems may repeat with AI workloads

• Why UX matters when operators need to trust the answer

• Where AI can help SRE and platform teams today

Links

Mat Ryer on LinkedIn: https://www.linkedin.com/in/matryer/

Mat Ryer on GitHub: https://github.com/matryer

Grafana Labs: https://grafana.com

Grafana Assistant: https://grafana.com/products/cloud/ai-assistant/

Grafana AI Observability: https://grafana.com/docs/grafana-cloud/machine-learning/ai-observability/

Grafana Adaptive Telemetry: https://grafana.com/products/cloud/adaptive-telemetry/

Grafana MCP server: https://github.com/grafana/mcp-grafana

OpenTelemetry: https://opentelemetry.io

Prometheus: https://prometheus.io

Grafana Loki: https://grafana.com/docs/loki/latest/

More episodes and show notes: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

GitHub API Enumeration, Grok Build CLI Data Exposure, AWS Security Hub Network Scanning, AI-Powered Patch Pressure, and Why Visibility Is Not Ownership18 Jul 202600:19:16

This week on Ship It Weekly: Datadog tracked coordinated GitHub API enumeration, xAI’s Grok Build CLI reportedly uploaded repo data without redaction, AWS Security Hub added Network Scanning and exposure impact analysis, and Microsoft says AI-powered vulnerability discovery is changing patch pressure.

The theme: visibility is not ownership. A GitHub API map does not revoke a token. An exposure finding does not close a port. A patch bulletin does not patch the fleet. And an AI coding tool reading your repo is still access.

Brian covers GitHub as a production surface, AI coding tool data boundaries, cloud exposure based on reachability and blast radius, and why patching needs to look more like production operations than spreadsheet theater.

Also, the Ship It Weekly shop is open at shop.tellerstech.com with Ship It Weekly t-shirt designs. Use coupon code SHIPTHESTORE for 20% off your order for the next few weeks.

Links

Datadog: Coordinated GitHub API enumeration https://securitylabs.datadoghq.com/articles/coordinated-github-api-enumeration/

The Verge: Grok Build CLI repository upload report https://www.theverge.com/ai-artificial-intelligence/965600/spacexai-grok-build-repository-upload

AWS Security Hub Network Scanning https://aws.amazon.com/about-aws/whats-new/2026/07/aws-security-hub-network-scanning/

AWS Security Hub impact analysis for exposure findings https://aws.amazon.com/about-aws/whats-new/2026/07/impact-analysis-aws-security-hub/

Microsoft: Windows vulnerability management and AI-powered discovery https://blogs.windows.com/windowsexperience/2026/07/09/evolving-windows-vulnerability-management-to-meet-the-speed-of-ai-powered-discovery/

SRE Weekly Issue 525 https://sreweekly.com/sre-weekly-issue-525/

HalluSquatting / hallucinated package risk https://www.endorlabs.com/learn/slopsquatting-when-ai-agents-hallucinate-malicious-packages

AWS Lambda Managed Instances for Java cold starts https://aws.amazon.com/blogs/compute/eliminating-java-cold-starts-with-aws-lambda-managed-instances/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W29/

Ship It Weekly shop https://shop.tellerstech.com/

More episodes and full show notes https://shipitweekly.fm/

EKS Rollbacks, GitHub Actions Supply Chain Attacks, AI Agentjacking, CloudWatch Log Alarms, and Why Safety Nets Don’t Replace Ownership10 Jul 202600:19:13

This week on Ship It Weekly: Amazon EKS added Kubernetes version rollbacks, Novee Security published Cordyceps research on GitHub Actions supply chain risk, Tenet Security showed how fake telemetry can hijack AI coding agents, and Amazon CloudWatch added alarms directly from log queries.

The theme: safety nets are getting better, but the blast radius is getting wider. Rollback buttons, log alarms, zone-aware routing, secret scanning, and AI agent workflows all help, but they do not replace ownership.

Brian covers why EKS rollbacks are useful but not a substitute for real upgrade discipline, why GitHub Actions YAML is production code with credentials, how fake Sentry telemetry can become hostile agent context, and why easier log-based alarms can also mean easier pager noise.

In the lightning round: ECS Service Connect zone-aware routing, etcd 3.7, GitHub innersource advisories, secret scanning metadata improvements, and CloudWatch Application Signals service events.

Links

Amazon EKS Kubernetes version rollbacks https://aws.amazon.com/blogs/aws/upgrade-amazon-eks-clusters-with-confidence-using-kubernetes-version-rollbacks/

Novee Security: Cordyceps supply chain research https://novee.security/blog/cordyceps/

Tenet Security: Agentjacking through fake Sentry errors https://tenetsecurity.ai/blog/agentjacking-coding-agents-with-fake-sentry-errors/

Amazon CloudWatch log query alarms https://aws.amazon.com/about-aws/whats-new/2026/07/amazon-cloudwatch-log-alarms/

ECS Service Connect zone-aware routing https://aws.amazon.com/about-aws/whats-new/2026/07/ecs-service-connect-zone-aware/

etcd 3.7 announcement https://etcd.io/blog/2026/announcing-etcd-3.7/

GitHub innersource security advisories https://github.blog/changelog/2026-07-08-innersource-security-advisories-are-generally-available/

GitHub secret scanning extended metadata and multipart validation https://github.blog/changelog/2026-07-07-secret-scanning-extended-metadata-and-multipart-validation/

CloudWatch Application Signals service events https://aws.amazon.com/about-aws/whats-new/2026/06/cloudwatch-service-events/

Our Links

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W28/

More episodes and full show notes https://shipitweekly.fm/

Ship It Conversations: Evan Phoenix of Miren on Deployment Pain, Terraform, Waypoint, and Better Defaults for Small Teams06 Jul 202600:41:52

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Evan Phoenix of Miren about why deployment is still painful, what teams keep getting wrong when they try to simplify it, and why small teams may need better defaults more than more platform knobs.

Evan is the CEO of Miren. He previously worked on Terraform Enterprise and Waypoint at HashiCorp, and he also built Puma and Rubinius.

We talk about deployment as the “final boss” of software delivery. Not because teams do not know how to ship code, but because deployment is where everything collides: runtimes, registries, secrets, networking, cloud services, databases, rollbacks, and the internal platform nobody wants to touch anymore.

A big theme is opinionated tooling. Engineers often say they want flexibility, but many teams are really asking for good defaults, a clear happy path, and fewer decisions to own.

We also get into Terraform Enterprise, Terraform Cloud, Terragrunt, OpenTofu, Waypoint, Kubernetes, Heroku, ECS, container registries, and how AI changes the deployment conversation. AI can generate infrastructure code, but when that setup breaks, someone still has to understand it and be on call for it.

Highlights

• Why deployment is still painful after years of platforms and abstractions

• What Evan learned from Terraform Enterprise and Waypoint

• Why Terraform structure, state, modules, and repo layout remain hard

• Why OpenTofu gained traction beyond the Terraform licensing change

• Why Kubernetes can be too much surface area for some teams

• What small teams actually need from deployment tooling

• How AI changes infrastructure and deployment workflows

• Why generated infrastructure still needs ownership and accountability

Links

• Miren: https://miren.dev

• Miren on GitHub: https://github.com/mirendev

• Evan Phoenix: https://evanphx.dev

• Evan on Bluesky: https://bsky.app/profile/evanphx.dev

• Evan on Linkedin: https://www.linkedin.com/in/evanphoenix/

Things mentioned

• Terraform Enterprise: https://developer.hashicorp.com/terraform/enterprise

• Terraform Cloud: https://developer.hashicorp.com/terraform/cloud-docs

• Terragrunt: https://terragrunt.gruntwork.io

• OpenTofu: https://opentofu.org

• HashiCorp Waypoint: https://github.com/hashicorp/waypoint

• Knative: https://knative.dev

Our links

More episodes + show notes: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

Amazon Q CVEs, Hijacked npm and Go Packages, AWS WAF HTTP/2 Issues, Lambda MicroVMs, and Why Execution Is the Boundary Now03 Jul 202600:18:06

This week on Ship It Weekly: Amazon Q Developer and the AWS language servers had a pair of trust-boundary CVEs, JFrog found hijacked npm and Go packages using hidden VS Code tasks to run malware when a workspace opens, AWS WAF had HTTP/2 request-body inspection issues, and AWS introduced Lambda MicroVMs for running user-generated and AI-generated code in isolated sandboxes.

The bigger theme: execution is the boundary now. The repo, the IDE, the AI assistant, the WAF, and the sandbox all sit at the point where something gets to run, inspect, block, or decide. Before execution, trust is a policy. After execution, trust is a blast radius.

In the lightning round, Brian covers GitHub’s record advisory volume, Git 2.55, Valkey 9.1 on Amazon ElastiCache, and a quick Fable 5 callback now that Anthropic’s Fable 5 is back online.

Links

AWS security bulletin: Amazon Q / AWS language server CVEs https://aws.amazon.com/security/security-bulletins/2026-047-aws/

JFrog: Hijacked npm packages using VS Code tasks https://research.jfrog.com/post/hijacked-npm-vscode-tasks-blockchain/

AWS security bulletin: AWS WAF HTTP/2 inspection issues https://aws.amazon.com/security/security-bulletins/2026-048-aws/

AWS Lambda MicroVMs https://aws.amazon.com/blogs/aws/run-isolated-sandboxes-with-full-lifecycle-control-aws-lambda-introduces-microvms/

GitHub Advisory Database record volume https://github.blog/security/supply-chain-security/inside-the-advisory-database-and-what-happens-when-vulnerability-volume-breaks-records/

Git 2.55 highlights https://github.blog/open-source/git/highlights-from-git-2-55/

Amazon ElastiCache Valkey 9.1 https://aws.amazon.com/blogs/database/announcing-valkey-9-1-for-amazon-elasticache/

Claude Fable 5 and Mythos 5 model docs https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W27/

More episodes and full show notes https://shipitweekly.fm/

Ship It Conversations: Kat Traxler of Vectra AI on AI Security, the Zero-Day Clock, IAM, and Cloud Risk28 Jun 202600:42:33

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Kat Traxler of Vectra AI about AI security, the zero-day clock, IAM, cloud risk, AI-assisted bug hunting, and why the scariest future security problems may still start with the boring fundamentals teams already struggle with today.

Kat is a Principal Security Researcher at Vectra AI focused on abuse techniques and vulnerabilities in the public cloud, especially around the intersection of cloud security, AppSec, IAM, managed identities, and insecure-by-design flaws.

We talk about the current AI security mood, from the excitement around faster research and bug hunting to the fear that AI could shrink the window between vulnerability disclosure and exploitation. Kat explains the “San Francisco Consensus,” why the zero-day clock is getting so much attention, and why she thinks the facts may be real while some of the conclusions are overextended.

The bigger theme here is that AI is absolutely changing security work, but it does not erase the fundamentals. Attackers still take the lowest-friction path that works. For most teams, that still means credentials, IAM, misconfigurations, known vulnerabilities, and systems that were never threat-modeled as deeply as people assume.

Highlights

• Why AI security feels exciting and unsettling at the same time

• What the “San Francisco Consensus” means and why people are talking about the zero-day clock

• How AI may shrink the time between vulnerability disclosure and exploitation

• Why Kat is skeptical of the full “zero-day apocalypse” narrative

• Why credentials, IAM, misconfigurations, and known vulnerabilities still matter most for many teams

• How AI helps narrow the search space in bug hunting and security research

• Where AI is useful for code-level bugs, and where it still struggles with context and threat modeling

• Why human expertise still matters when using AI for writing, research, and cloud security analysis

• Why IAM remains hard because it sits at the intersection of people, access, and technology

• What insecure-by-design flaws are, and why AI may not solve those anytime soon

Kat / Vectra AI links

• Kat Traxler at Vectra AI: https://www.vectra.ai/about/author/kat-traxler

• Kat’s site: https://kattraxler.cloud/

• The San Francisco Consensus: https://kattraxler.cloud/the-san-francisco-consensus/

• Kat on X: https://x.com/NightmareJS

• Vectra AI: https://www.vectra.ai/

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

containerd CRI Vulnerabilities, Datadog PostgreSQL HA on Kubernetes, AWS DevOps Agent with Datadog MCP Server, EKS Control Plane Egress, and Why Users Feel the Wait26 Jun 202600:19:24

This week on Ship It Weekly: containerd disclosed a batch of CRI plugin vulnerabilities, Datadog tested PostgreSQL high availability on Kubernetes and found that failover is not useful if it cannot happen safely, AWS DevOps Agent and Datadog MCP Server moved AI incident response closer to real production workflows, and Amazon EKS added customer-routed control-plane egress.

The bigger theme: the control plane keeps getting wider. Runtimes, databases, incident agents, API-server egress, credentials, the cloud console, and object metadata are all becoming part of the production blast radius. And when something breaks, users do not experience your architecture diagram. They experience waiting.

In the lightning round, Brian covers GitHub self-service credential revocation for incident response, AWS Management Console Private Access without internet connectivity, Vercel Connect and short-lived agent credentials, and Amazon S3 annotations.

Links

containerd CRI plugin vulnerabilities / AWS security bulletin https://aws.amazon.com/security/security-bulletins/2026-046-aws/

Datadog: PostgreSQL high availability on Kubernetes https://www.datadoghq.com/blog/engineering/postgresql-ha-kubernetes/

AWS DevOps Agent and Datadog MCP Server https://aws.amazon.com/blogs/devops/production-ready-autonomous-incident-resolution-with-aws-devops-agent-now-ga-and-datadog-mcp-server/

Amazon EKS customer-routed control-plane egress https://aws.amazon.com/blogs/containers/amazon-eks-now-supports-control-plane-egress-through-your-vpc/

GitHub self-service credential revocation for incident response https://github.blog/changelog/2026-06-24-self-service-credential-revocation-for-incident-response/

AWS Management Console Private Access https://aws.amazon.com/about-aws/whats-new/2026/06/aws-management-console-private/

Vercel Connect https://vercel.com/blog/introducing-vercel-connect

Amazon S3 annotations https://aws.amazon.com/blogs/aws/amazon-s3-annotations-attach-rich-queryable-context-directly-to-your-objects/

Marc Brooker: Waiting, latency, MTTR, and the inspection paradox https://brooker.co.za/blog/2026/06/19/waiting.html

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W26/

More episodes and full show notes https://www.shipitweekly.fm

Ship It Conversations: Guardsquare’s Joel DeStefano on Mobile App Security, Runtime Protection, App Hardening, and Why Scanning Isn’t Enough21 Jun 202600:35:58

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Joel DeStefano from Guardsquare about mobile app security, why it is different from backend and cloud security, and why scanning alone is not enough once an app is shipped into the real world.

We talk about the shift in trust model that happens with mobile apps. In backend and cloud systems, teams usually have more control over the runtime, infrastructure, policies, and monitoring. With mobile, the app becomes a public artifact running on someone else’s device, in an environment you do not fully control.

The bigger theme here is that mobile security is not just “scan it before release.” Scanning matters, but teams also need to think about app hardening, obfuscation, runtime protection, monitoring, and whether the app connecting back to their APIs is genuine and uncompromised.

Highlights

• Why mobile changes the trust model compared to backend and cloud systems

• What DevOps, SRE, and platform teams should understand about mobile app risk

• Why scanning is useful, but not enough by itself

• The danger of assuming app store approval means an app is secure

• Why “we do not store sensitive data in the app” can be a misleading security argument

• How attackers can reverse engineer apps, inspect workflows, and learn how the app talks to backend APIs

• What code hardening and obfuscation actually help protect against

• Why runtime checks matter for rooted devices, compromised environments, debuggers, hooking frameworks, overlays, and accessibility abuse

• The difference between Android and iOS security assumptions

• Why the OS is not responsible for protecting your app’s business logic

• How mobile security should fit into CI/CD without destroying release velocity

• What should block a release versus what should become tracked risk

• Why testing, hardening, runtime protection, and monitoring should work together as one strategy

• How AI may speed up attackers without fundamentally changing the need for strong security fundamentals

• Joel’s advice for improving mobile security posture: start with the app’s critical workflows, backend interactions, and real business risk

Joel / Guardsquare links

• Guardsquare: https://hubs.ly/Q04fJgkJ0

• Guardsquare Blog: https://www.guardsquare.com/blog

OWASP mobile security links

• OWASP Mobile Application Security: https://owasp.org/www-project-mobile-app-security/

• OWASP MASVS: https://mas.owasp.org/MASVS/

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

PeopleSoft Zero-Day Exploited, npm v12 Install Script Changes, GitHub Agentic Tokens, Anthropic Model Risk, and Default Trust Breaking19 Jun 202600:22:27

This episode of Ship It Weekly is about default trust getting punished. Brian covers Oracle’s emergency PeopleSoft advisory for CVE-2026-35273, npm v12 changing install-script defaults, GitHub Agentic Workflows moving away from long-lived personal access tokens, and Anthropic disabling Fable 5 and Mythos 5 after a U.S. export-control directive. The common thread: legacy ERP systems, package installs, CI/CD agents, and AI models all become production risks when teams trust the default without checking what that trust can actually do.

In the lightning round, Brian covers Tekton CloudEvents moving to a dedicated events controller, NVIDIA Triton Inference Server 26.04 changing inference defaults, AWS Nitro Isolation Engine bringing formal verification to Graviton5-based isolation, and Homebrew 6.0 adding explicit trust for third-party taps. The bigger theme: production does not care why you trusted the default. It only cares what that default was allowed to do.

The bigger theme: production does not care why you trusted the default. It only cares what that default was allowed to do.

Links

Oracle PeopleSoft CVE-2026-35273 advisory https://www.oracle.com/security-alerts/alert-cve-2026-35273.html

npm v12 breaking changes https://github.blog/changelog/2026-06-09-upcoming-breaking-changes-for-npm-v12/

GitHub Agentic Workflows no longer need PATs https://github.blog/changelog/2026-06-11-agentic-workflows-no-longer-need-a-personal-access-token/

Anthropic Fable 5 / Mythos 5 access statement https://www.anthropic.com/news/fable-mythos-access

Tekton Pipelines releases https://github.com/tektoncd/pipeline/releases

NVIDIA Triton Inference Server 26.04 release notes https://docs.nvidia.com/deeplearning/triton-inference-server/release-notes/rel-26-04.html

AWS Nitro Isolation Engine https://aws.amazon.com/blogs/compute/aws-nitro-isolation-engine-formally-verifying-the-hypervisor-in-the-aws-nitro-system/

Homebrew 6.0.0 https://brew.sh/2026/06/11/homebrew-6.0.0/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W25/

More episodes and show notes https://shipitweekly.fm/

Ship It Conversations: Meta’s Francois Richard on AI Incident Response, SLOs, and Reliability at Scale16 Jun 202600:42:56

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Francois Richard, Engineering Director at Meta, about reliability at scale, how AI is changing production risk, what teams actually learn from incidents, and why recovery practice matters just as much as prevention.

We talk about the proactive and reactive sides of reliability, why SLOs should represent a promise to users instead of just another dashboard number, how incident reviews should drive real system improvements, and how teams can practice recovery before production forces the lesson on them.

The bigger theme here is that reliability is not just about avoiding failure. It is about knowing what happens when prevention fails. That means practicing regional failure, understanding overload behavior, improving incident response, using AI carefully during investigation, and making reliability targets match the actual lifecycle and importance of the system.

Highlights

• Why reliability work starts with both prevention and recovery

• The difference between reactive incident response and proactive reliability engineering

• How Meta thinks about disaster recovery testing and regional failure practice

• Why an SLO should be treated like a promise to users, not just a dashboard metric

• How SLO trends help teams decide when to invest more in reliability or take more product risk

• What engineers actually learn during the “pressure cooker” of an incident

• Why incident reviews should produce follow-up work, not just a nicer explanation of what broke

• The difference between finding the cause of an incident and improving the system

• Where AI agents can help with incident investigation, telemetry, metrics, and query building

• Why AI-generated code can increase change volume while reducing human context

• How faster code generation changes the kinds of reliability problems teams should expect

• Why recovery practice matters, especially for region loss, traffic spikes, overload, and restart behavior

• What smaller DevOps and SRE teams can learn from Meta-scale reliability patterns

• Why not every system needs six nines, especially early in a product lifecycle

• How to think about reliability investment based on user promise, product maturity, and operational risk

• Why At Scale Systems & Reliability is focused on the infrastructure behind AI and the use of AI to operate large-scale systems

Francois’ links

• LinkedIn: https://www.linkedin.com/in/francoisrichard/

At Scale links

• Systems & Reliability 2026: https://bit.ly/4xd2FdG

• At Scale Conferences: https://atscaleconference.com/

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

Coinbase Outage, Meta AI Account Recovery, AWS AgentCore Code Injection, Apigee Tenant Isolation, and the Glue That Breaks Production12 Jun 202600:23:11

This episode of Ship It Weekly is about the hidden glue holding production together.

Brian covers Coinbase’s May 7 outage postmortem, where an AWS us-east-1 cooling failure exposed the difference between being “multi-AZ” on paper and actually being able to recover when stateful, low-latency systems are tied to a failed zone.

Then he looks at Meta’s AI-assisted Instagram support issue and why account recovery is identity infrastructure, not just customer support. If AI can influence password resets, email changes, MFA resets, or account ownership flows, that workflow needs to be treated like a production control plane.

The episode also covers AWS AgentCore CLI CVE-2026-11393, where collaborator metadata could break out into generated Python code during agent import, and an Apigee cross-tenant issue from Google’s Apigee security bulletins that shows why tenant isolation has to be tested beyond the obvious happy path.

Links

Coinbase May 7 outage postmortem https://www.coinbase.com/blog/a-postmortem-of-our-may-7-2026-outage

Meta AI support / Instagram account recovery reporting https://www.theverge.com/tech/945658/meta-ai-support-chatbot-exploit-instagram-accounts

AWS AgentCore CLI CVE-2026-11393 https://aws.amazon.com/security/security-bulletins/2026-040-aws/

AgentCore CLI GitHub advisory https://github.com/aws/agentcore-cli/security/advisories/GHSA-m4x6-gwgp-4pm7

Google Apigee security bulletins https://docs.cloud.google.com/apigee/docs/security-bulletins/security-bulletins

Cloudflare real-time threat intel WAF rules https://blog.cloudflare.com/realtime-threat-intel-waf-rules/

AWS Lambda tenant isolation with event source mappings https://aws.amazon.com/blogs/compute/integrating-event-source-mappings-with-aws-lambda-tenant-isolation-mode/

Amazon OpenSearch Serverless next generation https://aws.amazon.com/about-aws/whats-new/2026/05/amazon-opensearch-serverless-next-generation-generally-available/

GitHub Enterprise Managed Users IP allow list coverage https://github.blog/changelog/2026-06-08-ip-allow-list-coverage-for-emu-namespaces-in-general-availability/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief-news/2026-W24/

More episodes and show notes https://shipitweekly.fm/

Kiro CLI Approval Bypass, Amazon Braket Pickle Risk, AWS Org Logging, KEDA Upgrades, and Automation’s Hidden Boundaries05 Jun 202600:20:27

This episode of Ship It Weekly is about automation’s hidden boundaries. Brian covers Kiro CLI CVE-2026-9255, where piped stdin could act like user approval, Amazon Braket SDK CVE-2026-9291 and the very normal Python pickle risk hiding inside quantum job results, AWS Organizations finally emitting CloudTrail events when accounts join or leave an org, and KEDA updates that remind us autoscaling upgrades are production behavior changes.

The bigger thread this week is that automation does not remove boundaries. It moves them. Approval paths, trusted data, account membership, scaling signals, platform access, and AI-generated output all need clear ownership and visibility.

Brian also covers Kubernetes Dashboard being archived with Headlamp as the path forward, Google Cloud Remote MCP Server for AlloyDB, Apache Kafka 4.3.0, and Atlassian’s AI-native SDLC productivity claims.

Sponsored by @Scale: Systems & Reliability, happening June 25 at the Meydenbauer Center in Bellevue, Washington. Register at https://bit.ly/4xd2FdG

Links

Kiro CLI CVE-2026-9255 https://aws.amazon.com/security/security-bulletins/2026-035-aws/

Amazon Braket SDK CVE-2026-9291 https://aws.amazon.com/security/security-bulletins/2026-036-aws/

AWS Organizations CloudTrail account events https://aws.amazon.com/about-aws/whats-new/2026/05/aws-organizations-cloudtrail/

KEDA v2.20.0 release https://github.com/kedacore/keda/releases/tag/v2.20.0

KEDA v2.19.0 release https://github.com/kedacore/keda/releases/tag/v2.19.0

Kubernetes Dashboard archived / Headlamp path forward https://kubernetes.io/blog/2026/06/04/dashboard-archived-what-now/

Google Cloud Remote MCP Server for AlloyDB https://cloud.google.com/blog/products/databases/alloydb-remote-mcp-server-now-ga

Apache Kafka 4.3.0 https://www.confluent.io/blog/apache-kafka-4-3-release-announcement/

Atlassian AI-native SDLC productivity claims https://www.atlassian.com/blog/software-teams/ai-native-sdlc

This week’s On Call Brief https://www.tellerstech.com/on-call-brief/2026-W23/

More episodes and show notes https://shipitweekly.fm/

GitHub Supply Chain Attacks, Railway’s GCP Outage, Discord’s Voice Failure, AWS Retry Changes, and Trusted Tool Risk29 May 202600:23:47

This episode of Ship It Weekly is about trusted tools becoming production dependencies. Brian covers a rough GitHub supply chain week, including the compromised Nx Console VS Code extension tied to exposed GitHub internal repositories and the Megalodon campaign abusing GitHub Actions workflows across thousands of public repos.

The bigger thread this week is that the tools around production are increasingly part of production. Brian also covers Railway’s GCP account suspension outage, Discord’s voice outage during a Kubernetes migration, AWS changing SDK retry behavior, CVE-2026-9133 in the RabbitMQ AWS plugin, and a Reddit story about stolen AWS keys turning into a $14,000 Bedrock bill.

Brian also touches on OpenTelemetry graduating from the CNCF, Claude Code security risk, GitLab Secrets Manager, Google Cloud AI spend caps, and a Redshift Python driver RCE.

Full source list and extra links are available on this episode’s page at shipitweekly.fm.

Links

Nx Console compromise https://www.stepsecurity.io/blog/nx-console-vs-code-extension-compromised

Megalodon GitHub Actions attack https://www.stepsecurity.io/blog/megalodon-mass-github-actions-secret-exfiltration-across-5-500-public-repositories

Railway GCP outage https://blog.railway.com/p/incident-report-may-19-2026-gcp-account-outage

Discord voice outage https://discord.com/blog/behind-the-scenes-of-the-3-25-26-voice-outage

AWS SDK retry changes https://aws.amazon.com/blogs/developer/announcing-updated-retry-behavior-for-aws-sdks-and-tools/

RabbitMQ AWS plugin CVE-2026-9133 https://aws.amazon.com/security/security-bulletins/2026-034-aws/

AWS Bedrock cost spike Reddit thread https://www.reddit.com/r/aws/comments/1tm3ydo/aws_bedrock_cost_spike_14000_usd/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief/2026-W22/

More episodes and show notes https://shipitweekly.fm/

Ship It Conversations: Jake Warner on Cycle.io, Bare Metal’s Comeback, and Why Private Cloud Is Getting Interesting Again26 May 202600:36:06

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Jake Warner, founder and CEO of Cycle.io, about private cloud, bare metal, Kubernetes fatigue, and why some teams are rethinking how much infrastructure complexity they actually want to carry.

We talk about why bare metal and private cloud are getting interesting again, especially around cost, performance, data sovereignty, compliance, and platform ownership. Jake explains how Cycle approaches infrastructure as a pool of resources, why he thinks in terms of “environments as code” instead of traditional infrastructure as code, and how teams can run containers and VMs together across bare metal, cloud, and hybrid environments.

The bigger theme here is that this is not really a “cloud versus bare metal” conversation. It is about choosing the right level of abstraction. Sometimes Kubernetes is the right answer. Sometimes managed cloud services make sense. And sometimes teams just need a more opinionated platform that lets developers ship without requiring a large DevOps army to keep everything running.

Highlights

• Why some teams are moving back toward private cloud and bare metal

• The role of cost, data sovereignty, compliance, and performance in infrastructure decisions

• Why bare metal does not have to mean going back to old-school racking and stacking pain

• How Cycle turns raw compute into a private cloud-style resource pool

• Why Jake thinks about “environments as code” instead of only infrastructure as code

• What “no DevOps army required” means in practice for engineering-heavy teams

• Why some companies need VMs and containers running together on the same platform

• Where Kubernetes still makes sense, especially for highly customized infrastructure needs

• Why opinionated platforms can be valuable when teams want fewer knobs and better defaults

• Active-active thinking, failover risk, and why application-level replication often matters more than platform-level storage magic

• Why bandwidth, performance density, and predictable pricing can make bare metal attractive again

• The weird continued gravity of AWS us-east-1, even for teams trying to move workloads elsewhere

• How AI workloads, GPUs, and hype cycles fit into the private cloud and platform conversation

• Jake’s advice for modernizing hybrid or on-prem infrastructure: containerize first, then look hard at your dependencies

Jake’s links

• Cycle.io: https://cycle.io/

• Cycle Slack community: https://slack.cycle.io/

• Jake Warner on LinkedIn: https://www.linkedin.com/in/jakewarner/

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

CISA’s GitHub Leak, AI Root Cause Analysis, Copilot Agents, Claude Code in CI/CD, and Kubernetes Seccomp Risk22 May 202600:22:23

This episode of Ship It Weekly is about secrets, agents, risky defaults, and follow-up work that never gets done. Brian covers the CISA contractor GitHub leak involving AWS keys, internal docs, Terraform, Kubernetes, Argo CD, and CI/CD context, plus AWS DevOps Agent doing automated RCA across Datadog, Elasticsearch, CloudTrail, and EKS.

Brian also covers MS Copilot Studio computer-using agents, Claude Code in Bitbucket Agentic Pipelines, CVE-2026-46333 and Kubernetes seccomp defaults, GitHub OIDC for Dependabot, Java pods getting OOMKilled, LLM-generated SQL that can be wrong but still run, and why postmortem action items die without ownership.

Sponsored by Guardsquare https://hubs.ly/Q04fJgkJ0

Links

CISA GitHub leak https://blog.gitguardian.com/how-we-got-a-cisa-github-leak-taken-down-in-26-hours/

AWS DevOps Agent RCA https://aws.amazon.com/blogs/devops/automate-root-cause-analysis-across-datadog-and-elasticsearch-with-aws-devops-agent/

Microsoft Copilot Studio computer-using agents https://techcommunity.microsoft.com/blog/copilot-studio-blog/computer-using-agents-in-microsoft-copilot-studio-are-now-generally-available/4519427

Atlassian Agentic Pipelines with Claude Code https://support.atlassian.com/bitbucket-cloud/docs/agentic-pipelines/

CVE-2026-46333 https://nvd.nist.gov/vuln/detail/CVE-2026-46333

Kubernetes seccomp https://kubernetes.io/docs/reference/node/seccomp/

GitHub OIDC for Dependabot and code scanning https://github.blog/changelog/2026-05-19-expanded-oidc-support-for-dependabot-and-code-scanning/

Java pods OOMKilled in Kubernetes https://dzone.com/articles/java-pod-oomkill-kubernetes

LLM-generated SQL risks https://readyset.io/blog/why-llms-write-incorrect-sql-and-what-that-means-for-your-database

Postmortem action items https://incident.io/blog/why-do-post-mortem-action-items-fail-how-to-make-incident-follow-ups-actually-get-done

On Call Brief https://www.tellerstech.com/on-call-brief/2026-W21/

More episodes + show notes https://shipitweekly.fm/

AI Agents Get API Access and Identity: GitHub Copilot Cloud Agents, MCP Auth, Ansible Automation, OpenAI Daybreak, and the New Production Risk14 May 202600:23:21

This episode of Ship It Weekly is about AI agents moving from helpful coding assistants into real operational actors. Brian covers GitHub making Copilot cloud agent tasks available through a REST API, Auth0 bringing authentication and authorization to MCP servers, Red Hat positioning Ansible as a trusted execution layer for agentic IT operations, and OpenAI Daybreak pushing AI deeper into security research and remediation.

The bigger thread this week is authority: what these agents can reach, what they can change, who approved the action, and who owns the outcome when something breaks.

Brian also covers Discord’s ScyllaDB automation work, AWS GuardDuty crypto mining detection, queues and back pressure, and a Datadog PostgreSQL case where an index scan was still painfully slow.

Sponsored by Guardsquare https://hubs.ly/Q04fJgkJ0

Links

GitHub Copilot cloud agent tasks via REST API https://github.blog/changelog/2026-05-13-start-copilot-cloud-agent-tasks-via-the-rest-api/

GitHub REST API endpoints for agent tasks https://docs.github.com/en/rest/agent-tasks/agent-tasks

Auth0 Auth for MCP is now generally available https://auth0.com/blog/auth0-auth-for-mcp-servers-generally-available/

Red Hat on Ansible as the execution layer for agentic IT https://www.redhat.com/en/about/press-releases/red-hat-establishes-ansible-automation-platform-trusted-execution-layer-it-operations-agentic-era

OpenAI Daybreak https://openai.com/daybreak/

Discord automates ScyllaDB clusters at scale https://discord.com/blog/how-discord-automates-scylladb-clusters-at-scale

AWS GuardDuty crypto mining detection and prevention https://aws.amazon.com/blogs/security/detecting-and-preventing-crypto-mining-in-your-aws-environment/

Queues do not absorb load, they delay failure https://dzone.com/articles/queues-dont-absorb-load-they-delay-bankruptcy

Datadog on inefficient PostgreSQL index scans https://www.datadoghq.com/blog/detect-inefficient-index-scans-with-dbm/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief/2026-W20/

More episodes and show notes https://shipitweekly.fm/

Cursor Deletes PocketOS Prod DB, .de DNSSEC Outage, Bluesky Postmortem, Argo CD, and Copy Fail08 May 202600:21:57

This episode of Ship It Weekly is about modern reliability getting squeezed from both directions. Old-school failures still hit hard, like broken DNSSEC, kernel privilege escalation bugs, and GitOps behavior changes. But newer automation layers add a second kind of risk, where AI agents, machine identity, and cloud control planes can do real damage fast when authority is too broad. Brian covers the Cursor and PocketOS production database wipe, the .de DNSSEC outage and Cloudflare’s response, Bluesky’s April outage postmortem, Argo CD v3.1.16 reaching end of life plus the v3.4.1 behavior change, Linux kernel CVE-2026-31431 under active exploitation, and why Google Cloud Agent Identity and AWS MCP Server GA both point to agents becoming first-class infrastructure actors.

Sponsored by Guardsquare https://hubs.ly/Q04fJgkJ0

Links

Cursor / PocketOS production database wipe https://www.tellerstech.com/on-call-brief/2026-W19/

Cloudflare on the .de DNSSEC outage https://blog.cloudflare.com/de-tld-outage-dnssec/

Bluesky April 2026 outage postmortem https://pckt.blog/b/jcalabro/april-2026-outage-post-mortem-219ebg2

Argo CD releases: v3.1.16 final release and v3.4.1 behavior change https://github.com/argoproj/argo-cd/releases

Linux kernel CVE-2026-31431 https://nvd.nist.gov/vuln/detail/CVE-2026-31431

AWS bulletin for CVE-2026-31431 https://aws.amazon.com/security/security-bulletins/rss/2026-026-aws/

Google Cloud Agent Identity https://cloud.google.com/blog/products/identity-security/whats-new-in-iam-security-governance-and-runtime-defense

AWS MCP Server is now generally available https://aws.amazon.com/blogs/aws/the-aws-mcp-server-is-now-generally-available/

Cross-region disaster recovery for Amazon EKS using AWS Backup https://aws.amazon.com/blogs/containers/cross-region-disaster-recovery-for-amazon-eks-using-aws-backup/

Google Ads new data retention policy starting June 1, 2026 https://ads-developers.googleblog.com/2026/05/new-data-retention-policy-for-google.html

This week’s On Call Brief https://www.tellerstech.com/on-call-brief/2026-W19/

More episodes and show notes https://shipitweekly.fm/

Ship It Conversations: Gareth Kersey on IaCConf 2026, AI, and Corey Quinn’s Terraform Keynote05 May 202600:31:54

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

This episode is not sponsored. I wanted to cover IaCConf because the theme lines up closely with what Ship It Weekly focuses on: infrastructure, platform engineering, DevOps, SRE, and how teams are adapting to AI-driven change.

In this Ship It: Conversations episode, I talk with Gareth Kersey about IaCConf 2026, a free virtual conference focused on infrastructure as code, platform engineering, DevOps, SRE, and infrastructure operations. The conference is May 14th 2026.

The main theme is “keeping pace.” Not just keeping pace with new tools, but keeping pace with the speed of software delivery now that AI is changing how quickly application teams can write, ship, and change code.

We talk about what that means for the infrastructure teams underneath it all: the people responsible for Terraform, Kubernetes, GitOps, policies, secrets, cost, security, rollback paths, and making sure faster delivery does not turn into faster chaos.

Gareth walks through the IaCConf 2026 agenda, including Corey Quinn’s keynote, AI and Terraform sessions, platform engineering panels, Kubernetes and Argo CD talks, AI agents managing infrastructure as code, governance challenges, and the risk of 10x code velocity becoming 10x operational risk.

The bigger theme here is that AI is not just changing how code gets written. It is changing the pressure on the systems around delivery. Infrastructure as code, platform engineering, policy, and operational guardrails matter even more when the pace of change goes up.

Highlights

• What “keeping pace” means for infrastructure, DevOps, SRE, and platform teams

• Why faster application development can create more downstream operational pressure

• Corey Quinn’s keynote, “AI Speaks Terraform Like a Tourist”

• How AI-generated infrastructure changes create new governance and review challenges

• Why infrastructure as code still matters as AI agents and automation become more common

• Sessions covering Terraform, Kubernetes, Argo CD, GitOps, platform engineering, and AI-driven workflows

• The risk of 10x code velocity turning into 10x operational risk

• How platform teams can support faster developers without giving up safety or governance

• Why IaCConf includes panels, demos, technical talks, and practitioner stories instead of only tool-specific content

• How IaCConf has grown from its first event in 2025 into a broader infrastructure community

• Why the event is trying to stay community-focused instead of becoming just another vendor marketing conference

• The role of feedback, future spotlight events, in-person meetups, and possible community spaces around IaCConf

• Why registering still makes sense even if you cannot attend live, since sessions are available afterward

IaCConf links

• IaCConf 2026 registration page - https://www.iacconf.com/iacconf-2026

• IaCConf LinkedIn page - https://www.linkedin.com/showcase/iac-conf/

• IaCConf: https://www.iacconf.com/

• IaCConf is supported by Spacelift: https://spacelift.com

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

GitHub RCE, AI Agent Prompt Injection, and the New Reality: Your Developer Toolchain Is Production Now01 May 202600:25:08

This episode of Ship It Weekly is about the developer toolchain becoming part of production. Brian covers GitHub’s critical git push RCE, AI-assisted reverse engineering, prompt injection against AI agents in GitHub workflows, Elementary’s malicious CLI release, GitHub’s merge queue regression, Cal.com going closed source, and Copilot moving toward usage-based billing. Plus: MinIO’s repo archive, Ghostty leaving GitHub, Docker Hardened Images, and Azure DevOps security updates.

Links

GitHub git push RCE https://github.blog/security/securing-the-git-push-pipeline-responding-to-a-critical-remote-code-execution-vulnerability/

AI-assisted reverse engineering https://www.darkreading.com/application-security/reverse-engineering-ai-unearths-high-severity-github-bug

AI agents + GitHub Actions prompt injection https://www.theregister.com/2026/04/15/claude_gemini_copilot_agents_hijacked/

Elementary malicious CLI release https://www.elementary-data.com/post/security-incident-report-malicious-release-of-elementary-oss-python-cli-v0-23-3

GitHub merge queue regression https://github.blog/news-insights/company-news/an-update-on-github-availability/

Cal.com going closed source https://cal.com/blog/cal-com-goes-closed-source-why

GitHub Copilot billing https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/

MinIO archived repo https://github.com/minio/minio

Ghostty leaving GitHub https://mitchellh.com/writing/ghostty-leaving-github

Docker Hardened Images https://www.docker.com/blog/why-we-chose-the-harder-path-docker-hardened-images-one-year-later/

Azure DevOps security updates https://devblogs.microsoft.com/devops/one-click-security-scanning-and-org-wide-alert-triage-come-to-advanced-security/

On Call Brief https://oncallbrief.com/

More episodes https://shipitweekly.fm/

Kubernetes 1.36, Gateway API v1.5, AWS Copilot End of Support, and Cloudflare Non-Human Identities24 Apr 202600:20:24

This episode of Ship It Weekly is about platforms getting sharper about defaults, ownership, and the old paths they are no longer willing to quietly carry forever. Brian covers Kubernetes 1.36 and why it feels more like a cleanup-and-maturity release than a flashy feature dump, Gateway API v1.5 moving more networking behavior into the stable path, AWS Copilot CLI reaching end of support and what that means for teams still sitting on the older “easy” ECS workflow, Airbnb’s alert-development overhaul and why noisy or weak alerts are often a workflow problem long before they become an on-call problem, and Cloudflare’s push to treat scripts, agents, and third-party tools like real identities with real blast radius. He also hits the latest Azure DevOps Server patches and Google’s OTLP metrics support for Cloud Monitoring.

Links

Kubernetes v1.36 release https://kubernetes.io/blog/2026/04/22/kubernetes-v1-36-release/

Gateway API v1.5 https://kubernetes.io/blog/2026/04/21/gateway-api-v1-5/

AWS Copilot CLI end of support https://aws.amazon.com/blogs/containers/announcing-the-end-of-support-for-the-aws-copilot-cli/

Airbnb on alert development https://medium.com/airbnb-engineering/it-wasnt-a-culture-problem-upleveling-alert-development-at-airbnb-01e2290eb0f5

Cloudflare on non-human identities, OAuth visibility, and scoped permissions https://blog.cloudflare.com/improved-developer-security/

Azure DevOps Server April patches https://devblogs.microsoft.com/devops/april-patches-for-azure-devops-server/

OTLP metrics for Google Cloud Monitoring https://cloud.google.com/blog/products/management-tools/otlp-opentelemetry-protocol-for-google-cloud-monitoring-metrics

Past episode where we talked about Cloudflare Mesh https://www.tellerstech.com/ship-it-weekly/aws-interconnect-ga-cloudflare-mesh-gitlab-19-eks-auto-mode-and-opentelemetry-config/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief/2026-W16/

On Call Brief: https://oncallbrief.com/

More episodes and show notes https://shipitweekly.fm/

Ship It Conversations: Stephane Moser on Pipedrive’s Jenkins-to-GitHub Actions Migration, Argo CD, and CI/CD at Scale19 Apr 202600:51:05

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Stephane Moser about Pipedrive’s move from Jenkins to GitHub Actions, building self-hosted runners on Kubernetes, shifting deployments toward GitOps with Argo CD, and what it actually takes to roll out a big CI/CD change across a large engineering org.

We talk about why Jenkins had become painful, from Groovy friction to noisy-neighbor problems on shared VMs, why GitHub Actions fit better, how reusable workflows and custom actions helped, why Argo CD beat out Flux for their use case, and how they had to build better observability and internal deployment visibility around GitHub as they scaled.

The bigger theme here is that this was not just a tooling swap. It was a product and platform migration. Isolation, repeatability, self-service, rollout strategy, and observability mattered just as much as the actual CI/CD tools.

Highlights

• Why Jenkins stopped working well for them: Groovy friction, shared VM contention, and poor predictability

• Replacing CodeShip pull request validation first as the low-blast-radius starting point

• Using Actions Runner Controller on Kubernetes with EKS and Karpenter for self-hosted runners

• Why reusable workflows and custom actions helped cut repetition across hundreds of services

• Choosing Argo CD over Flux, Argo Workflows, Tekton, and even a short Spinnaker attempt

• Moving from push-based deploys toward GitOps for better isolation and safer credentials handling

• Building internal observability because GitHub’s workflow visibility was not enough at their scale

• Dogfooding first, then rolling migration out in batches until teams could self-serve the move

• What broke when the new system actually worked too well: bot-driven deploy volume, queueing, and fairness

• The mobile side of the story: Mac minis, unstable runners, GitHub-hosted runners, and a very different migration path

• How AI sped up parts of the mobile migration and troubleshooting, without making the migration trivial

• Stephane’s advice for big CI/CD shifts: start small, reduce blast radius, and use your own platform first

Stephane’s links

• LinkedIn: https://www.linkedin.com/in/moserss/

• Talk video: https://www.youtube.com/watch?v=VrE1dh-1zEY

• Blog post Part 1: https://medium.com/pipedrive-engineering/so-long-jenkins-hello-github-actions-pipedrives-big-ci-cd-switch-03be29c75f63

• Blog post Part 2: https://medium.com/pipedrive-engineering/all-aboard-the-github-actions-express-pipedrives-big-ci-cd-switch-part-2-fcacf834afd2

• GitHub: https://github.com/moser-ss

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

AWS Interconnect GA, Cloudflare Mesh, GitLab 19, EKS Auto Mode, and OpenTelemetry Config17 Apr 202600:15:00

This episode of Ship It Weekly is about networking, ingress, and private access moving further up into the platform layer. Brian covers AWS Interconnect going generally available, Cloudflare Mesh, GitLab 19.0 breaking changes around Gateway API and bundled services, EKS Auto Mode networking, and OpenTelemetry declarative config reaching stability. He also hits containerd security patches, GitHub’s new Code Security risk assessment, and AWS guidance on securing AI agents with MCP. (Amazon Web Services, Inc.)

Links

AWS Interconnect GA and last mile connectivity https://aws.amazon.com/blogs/aws/aws-interconnect-is-now-generally-available-with-a-new-option-to-simplify-last-mile-connectivity/

Cloudflare Mesh https://blog.cloudflare.com/mesh/

GitLab 19.0 breaking changes https://about.gitlab.com/blog/a-guide-to-the-breaking-changes-in-gitlab-19-0/

EKS Auto Mode networking https://aws.amazon.com/blogs/containers/navigating-enterprise-networking-challenges-with-amazon-eks-auto-mode/

OpenTelemetry declarative config reaches stability https://opentelemetry.io/blog/2026/stable-declarative-config/

containerd security releases https://github.com/containerd/containerd/releases

GitHub Code Security risk assessment for organizations https://github.blog/changelog/2026-04-08-code-security-risk-assessment-available-for-organizations/

AWS secure AI agent access patterns using MCP https://aws.amazon.com/blogs/security/secure-ai-agent-access-patterns-to-aws-resources-using-model-context-protocol/

This week’s On Call Brief https://www.tellerstech.com/on-call-brief/2026-W16/

More episodes and show notes https://shipitweekly.fm/

Special: Claude Mythos Preview and Project Glasswing: AI Exploit Discovery, Zero-Day Risk, Business Fallout, and What It Means for DevOps, Cloud, and Platform Security16 Apr 202600:16:28

In this Ship It Weekly special, Brian breaks down Claude Mythos Preview and Project Glasswing, and why this story matters beyond normal AI launch hype.

Anthropic is treating Mythos like a real security inflection point, not just a better coding model. Project Glasswing is their coordinated effort to get early access into the hands of defenders, critical software maintainers, and major infrastructure organizations before similar capability becomes more broadly available. If OpenClaw was about agents becoming a new control plane, this episode is about what happens when finding ways into messy environments and control planes starts getting faster too.

We walk through the practical angle for DevOps, cloud, platform, and infra teams: exploit timelines may be compressing, platform debt becomes attacker leverage, and the boring work most orgs treat like cleanup suddenly looks a lot more like frontline security work. We also zoom out to the business side, including why banks, regulators, and government officials are already paying attention.

Chapters

  • Why This Episode Exists
  • OpenClaw Callback
  • What Actually Happened
  • Don’t Get Gullible, Don’t Get Lazy
  • What Changes If This Is Even Half True
  • Why Business People Should Care
  • What This Means for DevOps, Cloud, and Platform
  • Boring Work Just Got Promoted
  • The Uncomfortable Takeaway
  • What I’d Do Right Now

Links from this episode

Claude Mythos Preview

https://red.anthropic.com/2026/mythos-preview/

Project Glasswing

https://www.anthropic.com/project/glasswing

AI cyber threats: open letter to business leaders

https://www.gov.uk/government/publications/ai-cyber-threats-open-letter-to-business-leaders/ai-cyber-threats-open-letter-to-business-leaders-html

AI-boosted hacks with Anthropic’s Mythos could have dire consequences for banks

https://www.reuters.com/legal/litigation/ai-boosted-hacks-with-anthropics-mythos-could-have-dire-consequences-banks-2026-04-13/

ECB to quiz bankers about risks of Anthropic's new AI model, source says

https://www.reuters.com/world/ecb-warn-bankers-about-new-anthropic-model-risks-source-says-2026-04-15/

Related episode: OpenClaw special

https://www.tellerstech.com/ship-it-weekly/special-openclaw-security-timeline-and-fallout-cve-2026-25253-one-click-token-leak-malicious-clawhub-skills-exposed-agent-control-panels-and-why-local-ai-agents-are-a-new-devops-sre-control-plane/

Amazon S3 Files, Malicious npm Plugins, Trivy Fallout, and Kubernetes’ Gateway Shift10 Apr 202600:15:04

This episode of Ship It Weekly is about the interface layer becoming the story. Brian covers Amazon S3 Files and why it feels more like a managed filesystem layer in front of S3 than “S3 is EFS now,” including how it relates to the old s3fs and FUSE-style approach. He also digs into 36 malicious npm packages posing as Strapi plugins, the uglier follow-on to the Trivy incident he discussed previously, Kubernetes Ingress2Gateway 1.0 and the push toward Gateway API, and Kubernetes Agent Sandbox as a sign that newer AI-style workloads are starting to reshape the platform itself.

Links

Amazon S3 Files

https://aws.amazon.com/blogs/aws/launching-s3-files-making-s3-buckets-accessible-as-file-systems/

Malicious npm packages posing as Strapi plugins

https://thehackernews.com/2026/04/36-malicious-npm-packages-exploited.html

Trivy follow-on incident discussion

https://github.com/aquasecurity/trivy/discussions/10425

RoseSecurity on Trivy / typosquatting angle

https://rosesecurity.dev/2026/03/20/typosquatting-trivy.html

Earlier episode covering the first Trivy incident

https://www.tellerstech.com/ship-it-weekly/aws-bahrain-uae-data-center-issues-amid-iran-strikes-argocd-vs-flux-gitops-failures-github-actions-hackerbot-claw-attacks-trivy-roguepilot-codespaces-prompt-injection-block-ai-remake/

Kubernetes Ingress2Gateway 1.0

https://kubernetes.io/blog/2026/03/20/ingress2gateway-1-0-release/

Kubernetes Agent Sandbox

https://kubernetes.io/blog/2026/03/20/running-agents-on-kubernetes-with-agent-sandbox/

Fortinet FortiClient EMS emergency patch

https://www.fortiguard.com/psirt/FG-IR-26-099

Karpathy post

https://x.com/karpathy/status/2036487306585268612

ProofShot

https://github.com/AmElmo/proofshot

More episodes and show notes

https://shipitweekly.fm

On Call Briefs

https://oncallbrief.com

Ship It Conversations: David Tuite on Backstage, Internal Developer Portals, and the Shift to AI Agents06 Apr 202600:33:55

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with David Chute, founder and CEO of Roadie, about internal developer portals, Backstage, automation, and how IDPs may evolve as AI agents become more common in engineering workflows.

We talk about the difference between a platform and a portal, the three common problems IDPs usually try to solve, why discoverability tends to be the first pain teams feel, and why a lot of orgs should start with automation before trying to perfect a service catalog. We also get into self-hosted Backstage vs managed options, and how teams should think about adoption, data models, and time to value.

The bigger theme is the one I found most interesting: IDPs may be shifting away from dashboard-heavy “single pane of glass” thinking and toward becoming context layers for workflows, terminals, and eventually agents.

Highlights

• The difference between an internal developer platform and an internal developer portal

• The three common IDP problem areas: discoverability, automation, and guardrails

• Why discoverability is usually the first pain teams feel

• Why adoption is often more of a human problem than a technical one

• Catalog completeness vs team ownership

• Why a lot of teams should start with automation first

• Self-hosted Backstage vs SaaS tradeoffs: extensibility, control, lock-in, and time to value

• Why IDPs may move from dashboards to context delivery for humans and agents

• Why AI helps teams build faster, but does not solve the problem of building the right thing

• David’s advice for platform and DevEx teams: talk to your internal users first

David’s links

• LinkedIn: https://www.linkedin.com/in/davidtuite/

Roadie / Backstage

• Roadie: https://roadie.io/

• Backstage: https://backstage.io/

Stuff mentioned

• Workday

• Backstage

• GitHub

• GitLab

• Bitbucket

• Azure DevOps

• Argo CD

• LaunchDarkly

• CircleCI

• DORA metrics

• MCP-style context for agents

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

GitHub Actions Hardening, Airbnb Config Rollouts, Cloudflare Rust Restarts, ECS Managed Daemons, and Terraform Access Controls03 Apr 202600:13:54

This episode of Ship It Weekly is about the quiet platform work that keeps things safe before they break. Brian covers GitHub Actions hardening in Kubernetes-related repos, Airbnb’s safer config rollouts, Cloudflare’s zero-downtime Rust restarts, Amazon ECS Managed Daemons, and HCP Terraform access controls with IP allow lists and temporary AWS permission delegation.

Links

GitHub Actions security roadmap

https://github.blog/news-insights/product-news/whats-coming-to-our-github-actions-2026-security-roadmap/

Airbnb config rollouts

https://medium.com/airbnb-engineering/safeguarding-dynamic-configuration-changes-at-scale-5aca5222ed68

Cloudflare graceful restarts for Rust

https://blog.cloudflare.com/ecdysis-rust-graceful-restarts/

Amazon ECS Managed Daemons

https://aws.amazon.com/about-aws/whats-new/2026/04/amazon-ecs-managed-daemons/

HCP Terraform IP allow lists

https://www.hashicorp.com/blog/hcp-terraform-adds-ip-allow-list-for-terraform-resources

HCP Terraform AWS permission delegation

https://www.hashicorp.com/blog/aws-permission-delegation-now-generally-available-in-hcp-terraform

GitHub secret scanning updates

https://github.blog/changelog/2026-03-10-secret-scanning-pattern-updates-march-2026/

GitHub secret scanning for AI coding agents

https://github.blog/changelog/2026-03-31-secret-scanning-extends-to-ai-coding-agents-via-the-github-mcp-server/

Codespaces GA with data residency

https://github.blog/changelog/2026-04-01-codespaces-is-now-generally-available-for-github-enterprise-with-data-residency

Kubernetes v1.36 sneak peek

https://kubernetes.io/blog/2026/03/30/kubernetes-v1-36-sneak-peek/

GKE Inference Gateway

https://cloud.google.com/kubernetes-engine/docs/concepts/about-gke-inference-gateway

More episodes and show notes

https://shipitweekly.fm

On Call Briefs

https://oncallbrief.com

Hackerbot-Claw Grows, Xygeni Tag Poisoning, GitHub Search HA, Windows SID Failures, and AI Skills Supply Chain27 Mar 202600:15:25

This episode of Ship It Weekly is about the places where convenience quietly turns into trust.

Brian revisits the Trivy story by zooming out to the bigger hackerbot-claw GitHub Actions campaign, then gets into the Xygeni tag-poisoning compromise, GitHub’s search high availability rebuild for GitHub Enterprise Server, Windows Server 2025 surfacing duplicate SID problems in cloned images, and the agent-skills ecosystem replaying package supply chain history. Plus: a quick lightning round on GitHub pausing self-hosted runner minimum-version enforcement and March secret scanning updates.

Links

OpenSSF advisory on active GitHub Actions exploitation https://seclists.org/oss-sec/2026/q1/246

Xygeni action compromise via tag poisoning https://www.stepsecurity.io/blog/xygeni-action-compromised-c2-reverse-shell-backdoor-injected-via-tag-poisoning

GitHub Enterprise Server search high availability rebuild https://github.blog/engineering/architecture-optimization/how-we-rebuilt-the-search-architecture-for-high-availability-in-github-enterprise-server/

Microsoft on duplicate SIDs and nongeneralized Windows Server 2025 images https://learn.microsoft.com/en-us/troubleshoot/exchange/administration/exchange-server-issues-on-incorrect-windows-server-image

Socket on supply chain security for skills.sh https://socket.dev/blog/socket-brings-supply-chain-security-to-skills

Snyk ToxicSkills research https://snyk.io/blog/toxicskills-malicious-ai-agent-skills-clawhub/

GitHub self-hosted runner minimum version enforcement paused https://github.blog/changelog/2026-03-13-self-hosted-runner-minimum-version-enforcement-paused/

GitHub secret scanning pattern updates, March 2026 https://github.blog/changelog/2026-03-10-secret-scanning-pattern-updates-march-2026/

More episodes and show notes at https://shipitweekly.fm

On Call Briefs at https://oncallbrief.com

Ship It Conversations: Ang Chen on Project Vera, AI Cloud Emulation, and Safer Infrastructure Testing23 Mar 202600:24:23

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Ang Chen from the University of Michigan about Project Vera, a cloud emulator built to help teams test infrastructure changes more safely before they touch real cloud.

We talk about why testing against real cloud APIs is slow, expensive, and risky, how Vera works under tools like Terraform and CloudFormation, what “high fidelity” actually means, and where a tool like this could fit in local dev and CI/CD.

The bigger theme is one I think matters a lot: if AI is going to play a real role in cloud operations, it probably needs a sandbox first, not direct access to production.

Note

This interview was recorded on February 13, 2026. Since then, Vera’s public project materials have expanded the framing a bit further around multi-cloud support and safe environments for agent learning, so keep that in mind while listening.

Highlights

• Why real cloud testing still creates cost, delay, and risk

• How Vera emulates cloud behavior at the API layer

• Where this could help with Terraform, CloudFormation, and CI/CD workflows

• Why “useful enough to catch real mistakes” may matter more than perfect emulation

• The limits, tradeoffs, and fidelity questions that still need to be solved

• Why safe training grounds may matter before AI agents touch real infrastructure

Ang’s links

• LinkedIn: https://www.linkedin.com/in/ang-chen-8b877a17/

• University of Michigan profile: https://eecs.engin.umich.edu/people/chen-ang/

• Publications: https://web.eecs.umich.edu/~chenang/pubs.html

Project Vera

• Project site: https://project-vera.github.io/

• GitHub: https://github.com/project-vera/vera

• The quest for AI Agents as DevOps: https://project-vera.github.io/blogs/cloudagent/cloudagent/

• No More Manual Mocks: https://project-vera.github.io/blogs/cloudemu/cloudemu/

Stuff mentioned

• A Case for Learned Cloud Emulators: https://dl.acm.org/doi/10.1145/3718958.3754799

• Cloud Infrastructure Management in the Age of AI Agents: https://dl.acm.org/doi/abs/10.1145/3759441.3759443

• LocalStack: https://www.localstack.cloud/

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

McKinsey AI Flaw, Kafka Goes Diskless, Google Buys Wiz, AWS Copilot Ends, and AI Gateway on Kubernetes20 Mar 202600:14:56

This week on Ship It Weekly, Brian looks at what happens when new interfaces create old responsibilities.

McKinsey patched a vulnerability in its internal AI tool Lilli, Kafka contributors are pushing a diskless-topics model that rethinks durability and replication in cloud environments, and Google officially closed Wiz acquisition in one of the biggest cloud-security moves. Plus: AWS is sunsetting Copilot CLI, Kubernetes launches an AI Gateway Working Group.

Links

McKinsey statement on Lilli

https://www.mckinsey.com/about-us/media/statement-on-strengthening-safeguards-within-the-lilli-tool

Kafka diskless topics proposal

https://cwiki.apache.org/confluence/display/KAFKA/The%2BPath%2BForward%2Bfor%2BSaving%2BCross-AZ%2BReplication%2BCosts%2BKIPs

Google completes acquisition of Wiz

https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/wiz-acquisition/

AWS Copilot CLI end-of-support

https://aws.amazon.com/blogs/containers/announcing-the-end-of-support-for-the-aws-copilot-cli/

Kubernetes AI Gateway Working Group

https://kubernetes.io/blog/2026/03/09/announcing-ai-gateway-wg/

Amazon Bedrock observability for first-token latency and quota consumption

https://aws.amazon.com/about-aws/whats-new/2026/03/amazon-bedrock-observability-ttft-quota/

Cloudflare JSON responses and RFC 9457 support for 1xxx errors

https://developers.cloudflare.com/changelog/post/2026-03-11-json-rfc9457-responses-for-1xxx-errors/

Amazon S3 source-region information in server access logs

https://aws.amazon.com/about-aws/whats-new/2026/02/amazon-s3-source-region-information/

AWS Config adds 30 new resource types

https://aws.amazon.com/about-aws/whats-new/2026/03/aws-config-new-resource-types/

Amazon Bedrock AgentCore Runtime stateful MCP server features

https://aws.amazon.com/about-aws/whats-new/2026/03/amazon-bedrock-agentcore-runtime-stateful-mcp/

More episodes and show notes at

https://shipitweekly.fm

On Call Briefs at

https://oncallbrief.com

Meta Buys Moltbook, Block AI Layoffs Get Messier, Atlassian Cuts Jobs, and GitHub Explains the Outages13 Mar 202600:16:56

This week on Ship It Weekly, Brian covers five “AI meets reality” stories that every DevOps, SRE, security, and platform team can learn from.

Block’s AI layoff story is getting messier as follow-up reporting pushes back on the original framing, Meta bought Moltbook and brought more attention to the trust and security problems already showing up around AI-agent platforms, and Atlassian cut about 10% of its workforce while saying AI is changing the skills and roles it needs. Plus: GitHub gives one of the more honest outage breakdowns we’ve seen lately, Anthropic and Mozilla show a more grounded AI use case with Claude finding real Firefox bugs, and there’s a quick lightning round on Bedrock AgentCore policy, Dependabot for pre-commit hooks, and Cloudflare’s latest threat report.

Links

Block layoffs follow-up

https://www.theguardian.com/technology/2026/mar/08/block-ai-layoffs-jack-dorsey

Meta acquires Moltbook

https://www.theguardian.com/technology/2026/mar/10/meta-acquires-moltbook-ai-agent-social-network

Wiz on Moltbook exposure

https://www.wiz.io/blog/exposed-moltbook-database-reveals-millions-of-api-keys

Atlassian team update

https://www.atlassian.com/blog/announcements/atlassian-team-update-march-2026

GitHub availability issues write-up

https://github.blog/news-insights/company-news/addressing-githubs-recent-availability-issues-2/

Anthropic + Mozilla Firefox security

https://www.anthropic.com/news/mozilla-firefox-security

Anthropic labor market report

https://www.anthropic.com/research/labor-market-impacts

AWS Bedrock AgentCore Policy GA

https://aws.amazon.com/about-aws/whats-new/2026/03/policy-amazon-bedrock-agentcore-generally-available/

GitHub Dependabot support for pre-commit hooks

https://github.blog/changelog/2026-03-10-dependabot-now-supports-pre-commit-hooks/

Cloudflare 2026 Threat Report

https://blog.cloudflare.com/2026-threat-report/

More episodes and show notes at

https://shipitweekly.fm

On Call Briefs at:

https://oncallbrief.com

Ship It Conversations: Yvonne Young on Linux Foundations, Mentorship, and Getting Job Ready in Cloud09 Mar 202600:30:54

This is a guest conversation episode of Ship It Weekly (separate from the weekly news recaps).

In this Ship It: Conversations episode I talk with Yvonne Young, a cloud and Linux mentor active in the CloudWhistler community. We talk about the real path into cloud and DevOps, why Linux still matters as a foundation, what “job ready” actually means, and why focus, consistency, and business thinking matter more than chasing every new tool.

Highlights

  • Linux fundamentals still matter because so much of cloud and infra work sits on top of Linux
  • What “job ready” really means: prepare for both technical and behavioral interviews, know the basics, and show how you learn when you don’t know something
  • Why so many juniors stall out by trying to learn everything instead of picking a direction
  • Why daily reps beat cramming: short, consistent practice keeps skills fresh better than marathon study sessions
  • How Yvonne thinks about certifications, including why hands-on certs like RHCSA stand out
  • Hands-on practice ideas: break things on purpose, troubleshoot, fix services, inspect ports, and use the help files
  • Why tools matter less than the business problem they solve
  • Using Vault as an example of solving real issues like secret sprawl, rotation, and centralized access
  • How to think about cloud learning: pick one provider, learn the concepts, and map your path to the kinds of companies you want to work for
  • Why mentorship and community matter, especially for juniors trying not to waste time or head in the wrong direction
  • What seniors can do better: better onboarding, real availability, and giving juniors an actual lifeline when they get stuck

Yvonne’s links

Stuff mentioned

More episodes + details: https://shipitweekly.fm

AWS Bahrain/UAE Data Center Issues Amid Iran Strikes, ArgoCD vs Flux GitOps Failures, GitHub Actions Hackerbot-Claw Attacks (Trivy), RoguePilot Codespaces Prompt Injection, Block “AI Remake” Layoffs, Claude Code Security07 Mar 202600:18:20

This week on Ship It Weekly, Brian looks at how the boundary of ops keeps expanding.

We cover AWS flagging issues in Bahrain/UAE amid Iran strikes, ArgoCD vs Flux and why ArgoCD can get stuck in failed sync states, GitHub Actions being exploited at scale (plus Trivy’s incident), RoguePilot prompt injection meeting real credentials in Codespaces, Block’s “AI remake” layoffs, and Anthropic’s Claude Code Security for defenders.

Lightning round: DeepSeek model access geopolitics, Vercel’s agentic security boundaries, a KEV CVE to patch, an MCP-atlassian SSRF-to-RCE chain, and Claude Cowork scheduled tasks.

Links

AWS Bahrain/UAE (Reuters) https://www.reuters.com/world/middle-east/amazon-cloud-unit-flags-issues-bahrain-uae-data-centers-amid-iran-strikes-2026-03-02/

ArgoCD to Flux https://hai.wxs.ro/migrations/argocd-to-flux/

GitHub Actions exploitation https://www.stepsecurity.io/blog/hackerbot-claw-github-actions-exploitation

Trivy incident https://github.com/aquasecurity/trivy/discussions/10265

RoguePilot https://thehackernews.com/2026/02/roguepilot-flaw-in-github-codespaces.html

Block layoffs (WSJ) https://www.wsj.com/business/jack-dorseys-block-to-lay-off-4-000-employees-in-ai-remake-28f0d869

Claude Code Security https://www.anthropic.com/news/claude-code-security

DeepSeek (Reuters) https://www.reuters.com/world/china/deepseek-withholds-latest-ai-model-us-chipmakers-including-nvidia-sources-say-2026-02-25/

Agentic boundaries https://vercel.com/blog/security-boundaries-in-agentic-architectures

CISA KEV https://www.cisa.gov/news-events/alerts/2026/03/03/cisa-adds-two-known-exploited-vulnerabilities-catalog

mcp-atlassian CVE https://arcticwolf.com/resources/blog-uk/cve-2026-27825-critical-unauthenticated-rce-and-ssrf-in-mcp-atlassian/

Claude Cowork tasks https://support.claude.com/en/articles/13854387-schedule-recurring-tasks-in-cowork

More: https://shipitweekly.fm

Cloudflare BYOIP BGP Withdrawals, Clerk’s Postgres Query-Plan Flip Outage, and AWS Kiro Permissions Lessons (Grafana Privesc + runc CVEs)27 Feb 202600:17:38

This week on Ship It Weekly, Brian looks at how the boundary of ops keeps expanding.

We cover AWS flagging issues in Bahrain/UAE amid Iran strikes, ArgoCD vs Flux and why ArgoCD can get stuck in failed sync states, GitHub Actions being exploited at scale (plus Trivy’s incident), RoguePilot prompt injection meeting real credentials in Codespaces, Block’s “AI remake” layoffs, and Anthropic’s Claude Code Security for defenders.

Lightning round: DeepSeek model access geopolitics, Vercel’s agentic security boundaries, a KEV CVE to patch, an MCP-atlassian SSRF-to-RCE chain, and Claude Cowork scheduled tasks.

Links

AWS Bahrain/UAE (Reuters) https://www.reuters.com/world/middle-east/amazon-cloud-unit-flags-issues-bahrain-uae-data-centers-amid-iran-strikes-2026-03-02/

ArgoCD to Flux https://hai.wxs.ro/migrations/argocd-to-flux/

GitHub Actions exploitation https://www.stepsecurity.io/blog/hackerbot-claw-github-actions-exploitation

Trivy incident https://github.com/aquasecurity/trivy/discussions/10265

RoguePilot https://thehackernews.com/2026/02/roguepilot-flaw-in-github-codespaces.html

Block layoffs (WSJ) https://www.wsj.com/business/jack-dorseys-block-to-lay-off-4-000-employees-in-ai-remake-28f0d869

Claude Code Security https://www.anthropic.com/news/claude-code-security

DeepSeek (Reuters) https://www.reuters.com/world/china/deepseek-withholds-latest-ai-model-us-chipmakers-including-nvidia-sources-say-2026-02-25/

Agentic boundaries https://vercel.com/blog/security-boundaries-in-agentic-architectures

CISA KEV https://www.cisa.gov/news-events/alerts/2026/03/03/cisa-adds-two-known-exploited-vulnerabilities-catalog

mcp-atlassian CVE https://arcticwolf.com/resources/blog-uk/cve-2026-27825-critical-unauthenticated-rce-and-ssrf-in-mcp-atlassian/

Claude Cowork tasks https://support.claude.com/en/articles/13854387-schedule-recurring-tasks-in-cowork

More: https://shipitweekly.fm

Ship It Conversations: Mike Lady on Day Two Readiness + Guardrails in the AI Era24 Feb 202600:34:38

This is a guest conversation episode of Ship It Weekly (separate from the weekly news recaps).

In this Ship It: Conversations episode I talk with Mike Lady (Senior DevOps Engineer, distributed systems) from Enterprise Vibe Code on YouTube. We talk day two readiness, guardrails/quality gates, and why shipping safely matters even more now that AI can generate code fast.

Highlights

  • Day 0 vs Day 1 vs Day 2 (launching vs operating and evolving safely)
  • What teams look like without guardrails (“hope is not a strategy”)
  • Why guardrails speed you up long-term (less firefighting, more predictable delivery)
  • Day-two audit checklist: source control/branches/PRs, branch protection, CI quality gates, secrets/config, staging→prod flow
  • AI agents: they’ll “lie, cheat, and steal” to satisfy the goal unless you gate them
  • Multi-model reviews (Claude/Gemini/Codex) as different perspectives
  • AI in prod: start read-only (logs/traces), then earn trust slowly

Mike’s links

Stuff mentioned

More episodes + details: https://shipitweekly.fm

Ship It Weekly – DevOps and SRE News for Engineers Who Run Production22 Feb 202600:00:53

Ship It Weekly is a DevOps and SRE news podcast for engineers who run real systems.

Every week I break down what actually matters in cloud, Kubernetes, CI/CD, infrastructure as code, and production reliability. No hype. No vendor spin. Just practical analysis from someone who’s been on call and shipped systems at scale.

This isn’t a tutorial show. It’s a signal filter.

I cover major industry shifts, security incidents, cloud provider changes, and tooling updates, then explain what they mean for platform teams and engineers operating in production.

If you work in DevOps, SRE, platform engineering, or cloud infrastructure and want context instead of clickbait, you’re in the right place.

New episodes weekly.

You can also find detailed write-ups at: https://shipitweekly.fm

And curated production-focused briefs at: https://oncallbrief.com

Subscribe, and let’s ship.

GitHub Agentic Workflows, Gentoo Leaves GitHub, Argo CD 3.3 Upgrade Gotcha, AWS Config Scope Creep20 Feb 202600:19:21

This week on Ship It Weekly, Brian hits five stories where the “defaults” are shifting under ops teams.

GitHub is bringing Agentic Workflows into Actions, Gentoo is migrating off GitHub to Codeberg, Argo CD upgrades are forcing Server-Side Apply in some paths, AWS Config quietly expanded coverage again, and EC2 nested virtualization is now possible on virtual instances.

Links

YouTube episodes https://www.youtube.com/watch?v=tuuLlo2rbI0&list=PLYLi5KINFnO7dVMbhsJQTKRFXfSSwPmuL&pp=sAgC

OnCallBrief https://oncallbrief.com

Teller’s Tech Substack https://tellerstech.substack.com/

GitHub Agentic Workflows (preview) https://github.blog/changelog/2026-02-13-github-agentic-workflows-are-now-in-technical-preview/

Gentoo moves to Codeberg https://www.theregister.com/2026/02/17/gentoo_moves_to_codeberg_amid/

Argo CD upgrade guide: 3.2 -> 3.3 (SSA) https://argo-cd.readthedocs.io/en/latest/operator-manual/upgrading/3.2-3.3/

AWS Config: 30 new resource types https://aws.amazon.com/about-aws/whats-new/2026/02/aws-config-new-resource-types

EC2 nested virtualization (virtual instances) https://aws.amazon.com/about-aws/whats-new/2026/02/amazon-ec2-nested-virtualization-on-virtual/

GitHub status page update https://github.blog/changelog/2026-02-13-updated-status-experience/

GitHub Actions: early Feb updates https://github.blog/changelog/2026-02-05-github-actions-early-february-2026-updates/

Runner min version enforcement extended https://github.blog/changelog/2026-02-05-github-actions-self-hosted-runner-minimum-version-enforcement-extended/

Open Build Service postmortem https://openbuildservice.org/2026/02/02/post-mortem/

Human story: AI SRE vs incident management https://surfingcomplexity.blog/2026/02/14/lots-of-ai-sre-no-ai-incident-management/

More episodes and show info on https://shipitweekly.fm

© My Podcast Data · Independent project · Data from Apple & Spotify