white decorative stripes white decorative stripes
About Pedab

PEDAB BLOG

Welcome to our world

Human Error Zero: Why the Biggest Threat to Network Uptime Isn't the Network

 

Ask any network operator what actually causes outages, and the honest answer is rarely a hardware failure or some exotic cyberattack. It's much simpler than that: someone touched the network, and something broke. Human error a bad configuration, a rushed change, software that wasn't quite ready, is still one of the leading causes of network downtime. The risk is so well understood that around major events like the Olympics or the year-end holidays, many service providers quietly agree on one rule: don't touch the network at all.

That instinct to freeze rather than improve is becoming harder to justify. Networks today face more pressure to change, and to change quickly, than at almost any point before. So the real question isn't whether networks will keep evolving ,they will, it's whether that evolution can happen without multiplying the very mistakes operators are trying to avoid.

Automation doesn't fix fragility — it speeds it up

There's a comforting assumption that automation solves the human-error problem: take the human out of the loop, and the errors disappear too. In practice, it's rarely that simple. Automating a network that already has reliability problems doesn't remove the mistakes, it just lets them happen faster, and at a bigger scale. Automation is only as trustworthy as the foundation underneath it. Get the hardware and software right first, and automation becomes something that reduces risk. Skip that step, and it just makes the same risk move quicker.

That's the logic behind what's increasingly being called a "human error zero" approach: build reliability into the foundation first, then layer automation and intelligence on top of it — not the other way around.

Why AI is turning up the pressure

Much of the current push to change networks comes down to AI. Industry groups like the GSMA have called "networks for AI" one of the defining trends of the year, covering both AI-optimized networking and the use of AI to help run networks themselves.

AI workloads roughly split into two types, and each puts a different kind of stress on the network. Training — the heavy lifting behind building large models — really doesn't tolerate lost data. If a GPU cluster loses packets mid-job, training has to roll back to its last checkpoint and start over, which is an expensive way to waste some of the priciest hardware in the building. Inference — the everyday experience of asking a model a question and getting an answer — cares less about the occasional lost packet and much more about speed and consistency. Every chatbot reply and real-time AI response depends on a network that behaves predictably, every time.

That split is quietly reshaping how data centers are built. The back-end connections between GPUs, long handled by a specialized protocol called InfiniBand, are increasingly shifting toward standard Ethernet — upgraded with new capabilities designed to make it just as reliable at that scale.

Moving compute closer to people

Inference has a geography problem: the closer the compute sits to the person using it, the faster it feels. That's nudging telecom operators toward smaller data centers spread out closer to where AI is actually being used by companies, everyday consumers, or increasingly, physical devices like robots, instead of relying purely on a handful of massive, centralized sites.

It's not only about speed. Smaller, distributed facilities are often easier to power than one giant site, and power has become a real constraint — several reports over the past year point to data center projects in the US being delayed or paused not for technical reasons, but simply because there wasn't enough power to run them. Distributed infrastructure also makes it easier for operators to make credible claims about data sovereignty, which matters more each year as high-profile public cloud outages keep reminding businesses just how much they depend on infrastructure they don't control.

Reliability isn't a slogan — it's an engineering decision

When researchers ask decision-makers what they actually want from data center infrastructure, reliability consistently tops the list — ahead of ease of integration, operations, or automation. That's not surprising: most people making these decisions have already lived through an outage they'd rather not repeat.

Getting there means treating reliability as two separate problems. First, the quality of the hardware and software itself — one rough but telling signal here is how many security flaws a vendor has had to patch over time, which hints at how much unplanned firefighting their customers end up doing. Second, the quality of day-to-day operations — and this is where the more interesting shift is happening: giving operators a safe way to test a change before it ever touches the live network.

Digital twins: testing changes before they can hurt you

More network platforms now come with a "digital twin", essentially a live, virtual copy of the network that lets operators try out a change safely before committing it for real. Pair that with transactional updates, where a batch of changes either all go through together or automatically rolls back if any part fails, and a network change stops being a leap of faith. It becomes something you can check, and undo. Real-time visibility adds to that: instead of working from a picture of the network that's a few minutes stale, operators see the moment something shifts, intentional or not, away from how the network is supposed to look.

Built on top of that foundation, AI agents are starting to earn their keep in a genuinely useful way. Not by replacing network engineers, but by doing the tedious first pass: sorting through alarms, cross-referencing log files from dozens of devices, and explaining, in plain language, what's most likely gone wrong. Work that used to eat an operator's whole afternoon can now take a few minutes. Worth being honest about the limits, though: an AI agent that tells you what's wrong is only useful to someone who already knows what "right" looks like. It speeds up the diagnosis, it doesn't replace the expertise needed to act on it.

Security can't be bolted on afterwards

As more companies route sensitive data and decisions through AI systems, that infrastructure becomes a more tempting target. DDoS attacks — floods of junk traffic designed to knock systems offline — have grown to a genuinely alarming scale, with recent record attacks measured in tens of terabits per second, several times more traffic than an entire mid-sized country would normally use. At the same time, the longer-term threat of quantum computing is pushing forward-thinking operators to adopt quantum-safe encryption now, years before quantum computers are expected to be capable of breaking today's codes — because data intercepted today could still be exposed once that day arrives.

From "don't touch it" to "trust it"

The thread running through all of this is a change in mindset. Networks used to be something you built once and then left alone. That approach simply doesn't hold up in a world where AI, edge computing, and security requirements are all pushing constant change. The alternative isn't to be reckless about it — it's to build enough trust into the foundation, through solid hardware and software, safe automation, real-time visibility, and the flexibility to work across vendors, that a Friday-afternoon change stops being something to dread.

Get that part right, and speed isn't a risk you're taking. It's just what naturally follows from doing things properly.

This kind of thinking was the whole point of a webinar session Nokia and Pedab ran together, "Human Error Zero: Networks that just work," where Nokia's Roland Thienpont walked through what reliable, change-friendly networking actually looks like in practice. If you'd rather see it explained live — with real demos and audience questions included — the full recording is below.


 

 

Curious how this would apply to your own network? Explore Nokia's data center networking solutions  and reach out to local Pedab representative. 

 

Inese Sustriņa

By: Inese Sustriņa

I am a digital marketing expert with over 10 years of experience in international B2B marketing, content creation, social media, and brand development. Having spent most of my career in the IT industry, I have developed a strong interest in technology and a natural drive to keep learning, allowing me to turn complex topics into clear, relevant, and engaging communication. I value authentic brand communication and focus on creating content that reflects each company’s unique identity while delivering meaningful business value.