Why Recovery Took Hours and Monitoring Missed the Warnings — What I Learned

When I entered the HQ technical department as a junior technician, I stepped into controlled chaos. Engineers stared at screens filled with scrolling text that looked like something from The Matrix, cables hanging all around, equipment scattered on every desk. People rushed in and out — dropping off broken equipment, picking up replacements, asking urgent questions.

I wondered how anyone could focus in this environment where interruptions came every few minutes.

I was in my third year of mechanical engineering university, barely knew any scripting or programming beyond playing with Linux at home. My entire network knowledge fit into a simple analogy: an IP address is like a zip code, the gateway is the postbox where you leave messages, and DNS translates human-readable addresses into IPs so computers can find each other.

I learned these concepts on my way to the interview for a technician position. The company was big, the pay was decent. I needed a job. Looking back, starting from the absolute bottom gave me something valuable: nowhere to go but up. And things moved fast. Weeks later, I was invited to join HQ and help the technicians with technical difficulties, serving as their point of contact. Months later, I was promoted to lead the team of field engineers.

Yes, me. A rookie with no formal studies in this field, leading a team where some members had university degrees in networking and telecommunications.

That’s a unique challenge — making people take you seriously when you have no credentials to justify the role. I couldn’t rely on diplomas or years of experience. I needed to earn their trust through actions, not titles. I needed to prove I was worth following by listening more than talking, by solving problems they couldn’t solve, and by making their jobs easier instead of harder.

Leadership without credentials means you serve first and lead second.

“This Is How We’ve Always Done It”

I quickly learned that things were done a certain way. Period. No questions asked.

This disturbed me. As a child, I’d argued with my grandfather: “I achieve the same results faster — why does the approach matter?” His response: “Because this is how we do it.” Here I was again, facing the same attitude. Inefficiencies everywhere. But as a junior, nobody takes you seriously.

The Department Everyone Ignored

I found others who were talked down to: the entire frontline support department. The people with direct customer contact. The ones who heard problems first.

When they came with ideas, they were dismissed. When they dared suggest something might be wrong with the network, they were put in their place. “The monitoring shows nothing, so the problem must be on the customer’s side, not ours.”

It was shocking.

Especially when minutes later, the network would tumble down. Those early warning signs from frontline support? Completely ignored. Next incident, same pattern. And again. And again.

Logistics Nightmare

When I took over managing the field engineers, my job became playing dispatch coordinator instead of engineering solutions. Which site should my team visit first? When time was tight, which customers would I leave without service until tomorrow?

Here’s how it worked: when we finally noticed an issue — either tickets had piled up or equipment already went dark in our monitoring — we dispatched a team to investigate. They’d reach the site and reboot equipment, check connections and LEDs for errors, then work their way down the line to find the problem.

But too many times, the equipment just needed reconfiguring or replacing. It was a logistics puzzle — deciding when teams should return to HQ to pick up equipment versus tackling another site.

Either way, downtime was high. We were always in recovery mode, constantly catching up, never getting ahead. Apparently, we needed more teams, more cars, more fuel — more resources.

This wasn’t the problem I should be solving. We should prevent downtimes and recover from them fast instead of throwing dice to decide who gets internet access today and who doesn’t.

I couldn’t ignore this.

The Single Point of Failure

Then I saw more: a bottleneck that was also a single point of failure. If we were already struggling in recovery mode, what happened when that single dependency failed? We’d be lost.

One PC. That’s all we had to connect and configure equipment. Everything — and everyone — depended on this one machine. Every piece of equipment that needed reconfiguring had to wait its turn. Every recovery process bottlenecked there.

If that PC died, our entire recovery capability died with it.

Initially, it seemed like more teams to reach more sites would help. But now it was clear: throwing more resources at a problem rarely solves it.

As Andy Benoit noted, progress often comes “not by deconstructing intricate complexities but by exploiting unrecognized simplicities.”

More teams would mean more complexity, more logistics coordination, more moving parts — but they’d all still end up waiting at that single PC for equipment to be configured one at a time.

Building the Bridge Nobody Built

Reducing downtime would require two things: better early warning of issues, and faster recovery when they happened. I started with the warning system.

The frontline support team held the key to early warnings — but nobody was listening to them. I built trust with them first, then created a simple dashboard where they could search for a customer and instantly see if there was a known issue in that area — something our automated monitoring couldn’t show them.

When multiple customers called about an area not marked as problematic, frontline informed me. I investigated, and if I found a real issue our monitoring missed, I marked the area as problematic.

Often, metrics tell you what happened. Patterns tell you what’s about to happen.

This created a feedback loop. They trusted the dashboard, started giving me feedback, and had the courage to speak up when something felt wrong. They weren’t dismissed anymore — they had a voice in the engineering structure.

Now I could dispatch my field engineers directly to real network issues instead of field technicians being sent on wild goose chases.

Breaking the Norm With Automation

Now came the real challenge: making my team of field engineers more efficient.

After insisting with management, we incrementally received laptops over the next few months until each team had one.

With laptops in hand, I broke the norm. Senior engineers insisted that single PC’s setup couldn’t be replicated on laptops — ‘it’s too complex, it won’t work’.

They had a point — the environment was complex. But I had to try, and to my surprise, it worked beautifully.

I discovered TCL-Expect — the best tool we had back then to adapt quickly to inputs. I started writing scripts that could identify equipment models when connected, then automatically reconfigure them.

Each device had a unique ID. I built a quick endpoint where the script could call HQ with that ID and retrieve the configuration.

For field engineers, it became simple: connect the equipment, run the script with the device ID, and watch it reconfigure automatically.

Every morning, I made sure each team had several “blank” devices in their car for different scenarios. If equipment couldn’t be reconfigured by the script, use a blank one and bring the broken one back in the evening.

This reduced trips to HQ substantially. My team gained independence. We switched from scrambling to catch up to recovering systems proactively.

That’s what engineering is about: building resilient systems with fast, explicit recovery paths in place so systems can be quickly restored when issues happen.

We’re talking about equipment powering entire apartment complexes or even entire blocks. A single failure could affect anything from a few customers to several thousand at once. The scale demanded automation — teaching every field engineer the ins and outs of different models of equipment from multiple manufacturers with different firmware versions simply wasn’t viable.

Downtime dropped from hours — sometimes overnight — to just minutes.

The Efficiency Paradox

The solution worked beautifully. Frontline identified patterns early. I could debug and mark areas as problematic. My field engineers could investigate and restore systems fast, dramatically improving customer experience.

This didn’t just save my team’s time. It saved frontline time, support team time, everyone’s time — reducing operational overhead across multiple departments.

And here’s the paradox: instead of needing more resources, this automation allowed me to reduce my team in half.

I kept the best engineers — the ones who were team players, who knew how to collaborate, who embraced more responsibility and appreciated the trust and freedom. These were people who wanted to improve things. The smaller, more focused team became faster, not slower. Combined with the early warning system from the frontline, we could now dispatch teams proactively to problem areas before issues escalated — getting ahead of failures instead of constantly reacting to customer complaints.

It reminds me of Max Verstappen throughout the 2025 season. Even when the results looked good, he refused to accept ‘good enough.’ That mindset isn’t about being at the top — it’s about how you get there. We were nowhere near the top — but as a united team, we stopped accepting ‘if it works, leave it like this.’

Permission to Fail

This wasn’t my designed path. As a former athlete whose sports career ended with an injury, I pivoted to mechanical engineering — I couldn’t afford a computer, so computer science seemed impossible. I only got one in my second year of university.

Yet there I was: learning networking on the way to an interview, months later leading experienced engineers, years later managing a county’s critical infrastructure as a self-taught Senior NOC engineer. (I once took down a quarter of the country’s internet — that’s another story.)

I was lucky to meet the right people who taught me leadership through their actions. The CTO once told me after I broke something: “The only way you’ll never break anything is if you don’t work at all. Now go fix it.” That permission to fail — and recover — shaped how I lead today.

Two decades later, I’m a Principal Software Engineer — Generalist. I’m not limited to a silo. I go where the constraint is, learn what’s needed, and do the work — whether that’s systems, architecture, code, or process.

Impostor syndrome hit me hard at every step. No credentials. No formal education in the field. No traditional path. Just a desire to keep learning and solve problems, focused on the results: systems running, problems solved, teams more efficient. I told myself: you’re doing it. Results matter more than diplomas.

I don’t know what the future holds, but I’m certain impostor syndrome will find me again if I take on completely new challenges. And when it does, I’ll remind myself: focus on the results, keep learning, and don’t be afraid to break things along the way and challenge the status quo.

Also published on LinkedIn