Asterisk Was Brilliant. Then We Fought It.
A short story about telephony, unexplained failure and knowing when infrastructure has consumed enough of your life
For a while, Asterisk was brilliant.
Alt Production Group was building out VoiP system. The IVR worked. Calls moved through the system. Things behaved approximately as telephones are generally expected to behave.
Excellent.
Then they didn't.
What followed was one of those infrastructure debugging experiences where you begin with:
“There will be a logical explanation for this.”
Then:
“That's strange.”
Then:
“Why are you doing that?”
And eventually:
“Right. Fuck this. We're replacing Asterisk.”
That last sentence is, occasionally, a legitimate architectural decision.
The problem with something that used to work
There is a particular frustration in debugging a system that has already demonstrated that it can function.
We weren't trying to make an entirely unsupported concept work.
It had worked.
Then behaviour changed, we started troubleshooting, and the amount of effort required to understand why it was no longer behaving in the same way became increasingly disproportionate to the thing we were actually trying to accomplish.
And eventually that becomes the engineering problem.
Not necessarily the original fault.
The operational burden created by diagnosing it.
Infrastructure-level components don't exist in isolation. Other systems and, ultimately, other people depend upon them.
That changes what “working” means.
A service that works perfectly until it fails and then requires an archaeological expedition through configuration, state and logs before anybody can understand what happened isn't necessarily resilient infrastructure.
Time-to-understanding matters
We talk a great deal about uptime.
We should probably talk more about time-to-understanding.
When something goes wrong:
Can the system explain what happened?
Can the operator isolate it?
Can components recover?
Can the architecture adapt?
Can we restore the service without requiring the person responsible for it to spend an entire day developing an increasingly personal vendetta against a PBX?
That last metric isn't currently an industry standard.
We think it deserves consideration.
The lesson we took from our Asterisk fight wasn't that Asterisk is universally bad. It has existed for decades, powers serious telephony deployments and clearly works for many organisations.
Our conclusion was much more specific.
If operating a component begins costing more engineering capacity than replacing that component, replacement becomes a rational engineering decision.
Infrastructure has to recover
That experience has influenced the way we think about the systems underneath Alt Production Group.
Infrastructure shouldn't merely be secure and performant when everything is healthy.
It needs to help operators understand what has happened when it isn't.
It should be capable of rapid diagnosis.
Where appropriate, it should adapt.
Recovery should be designed into the architecture rather than becoming something humans invent while the service is already broken.
Because when you're operating infrastructure, prolonged troubleshooting isn't an interesting technical puzzle.
It is downtime.
And there comes a point when the correct answer to:
“Why is Asterisk doing this?”
is no longer another three hours of troubleshooting.
Sometimes the answer is:
“It doesn't matter anymore. We're building something else.”
A first-person account from the Group’s founder, combining personal recollection and engineering judgement. For a factual correction, contact the press office with a non-sensitive outline.
Editorial & disclosure standards →