Theryxion
News and analysis from the world of production systems.
Multi-Region Failover Planning
May 3, 2026
Multi-region failover is mostly decided before the incident. The two questions that matter - how fresh does the standby data need to be, and who is allowed to press the button - sound managerial, but they drive every technical choice downstream from replication topology to health-check placement.
Data is the long pole. Application tiers scale horizontally and redeploy anywhere in minutes; a two-hundred-gigabyte database does not. Asynchronous replication buys you availability at the price of a recovery-point gap, and knowing your actual tolerance for that gap - in minutes, in euros - changes which fancy technologies are even admissible.
Managing Secrets Without Losing Sleep
May 14, 2026
There are exactly two ages of secrets management: 'we keep them in an encrypted file' and 'we were audited'. The distance between them is covered by rotation policies, access trails, and the gradual realisation that humans should read production credentials roughly never.…
When to Choose a Queue Over a Request
May 18, 2026
Traditional RPC calls remain popular due to straightforward causality: the client makes an invocation, waits for the response, and monitors latency directly. Asynchronous message queuing becomes essential when background tasks outlast active connections or when sudden volume surges threaten to overw…
The Operator's Guide to Load Testing
July 3, 2026
Benchmark simulations repeatedly fail to anticipate live incidents because synthetic request topologies overlook messy reality. Evenly distributed traffic aimed at single endpoints provides isolated micro-benchmarks. Real degradation occurs when synchronized retry floods slam backends after a moment…
A Practical Guide to API Rate Limiting
June 21, 2026
Most engineering teams acknowledge the necessity of rate limiting while neglecting its underlying architecture. Relying purely on client IP addresses collapses under carrier-grade NAT environments, where countless mobile subscribers route through shared gateways and end up throttled together as one …
More reading
- Zero-Downtime Deployments Without the Drama — Operations, July 9, 2026
- Practical Notes on Postgres Connection Pooling — Data, June 20, 2026
- Understanding TLS 1.3 Session Resumption — Security, July 3, 2026
About us
Our contributors have spent years on-call for large platforms. This site collects the playbooks, postmortems and reference material we wish someone had handed us earlier.