Theryxion

Theryxion

News and analysis from the world of production systems.

Operations

Multi-Region Failover Planning

May 3, 2026

Multi-region failover is mostly decided before the incident. The two questions that matter - how fresh does the standby data need to be, and who is allowed to press the button - sound managerial, but they drive every technical choice downstream from replication topology to health-check placement.

Data is the long pole. Application tiers scale horizontally and redeploy anywhere in minutes; a two-hundred-gigabyte database does not. Asynchronous replication buys you availability at the price of a recovery-point gap, and knowing your actual tolerance for that gap - in minutes, in euros - changes which fancy technologies are even admissible.

Continue reading →

Security

Managing Secrets Without Losing Sleep

May 14, 2026

There are exactly two ages of secrets management: 'we keep them in an encrypted file' and 'we were audited'. The distance between them is covered by rotation policies, access trails, and the gradual realisation that humans should read production credentials roughly never.…

Engineering

When to Choose a Queue Over a Request

May 18, 2026

Traditional RPC calls remain popular due to straightforward causality: the client makes an invocation, waits for the response, and monitors latency directly. Asynchronous message queuing becomes essential when background tasks outlast active connections or when sudden volume surges threaten to overw…

Engineering

The Operator's Guide to Load Testing

July 3, 2026

Benchmark simulations repeatedly fail to anticipate live incidents because synthetic request topologies overlook messy reality. Evenly distributed traffic aimed at single endpoints provides isolated micro-benchmarks. Real degradation occurs when synchronized retry floods slam backends after a moment…

Engineering

A Practical Guide to API Rate Limiting

June 21, 2026

Most engineering teams acknowledge the necessity of rate limiting while neglecting its underlying architecture. Relying purely on client IP addresses collapses under carrier-grade NAT environments, where countless mobile subscribers route through shared gateways and end up throttled together as one …

More reading

About us

Our contributors have spent years on-call for large platforms. This site collects the playbooks, postmortems and reference material we wish someone had handed us earlier.

More about the project →