Infrastructure reliability consultant

Vadim Alekseev

20+ years in IT operations. I built and led several network monitoring and operations teams, set up Problem Management in a bank, and ran independent audits followed by infrastructure optimization. Reading failure data and finding solutions is my profession - OpsLab is the tool: an exclusive statistical engine built for modern analysis and forecasting.

NOC / Monitoring Problem Management ITSM / ITIL Infrastructure operations Statistical analysis & forecasting

Track record

A NOC built from scratch, Problem Management in a bank, independent audits. 1999-2023.

See more →

My approach

The engine looks for the pattern. I propose the fix.

See more →

How we work together

How to start working together: a one-off review, an analysis inside your company, or bringing us in on a project.

See more →

What the module finds

Eight analysis methods with real numbers: priority, root causes, cycles, before/after, forecast, event links, MTBF/MTTR, industry norms.

See more →

Contacts

Email, LinkedIn, Telegram and the services deck as a file.

Get in touch →

Built reliability practices, not slide decks

2005-2019 · Rostelecom

Built the North-West Russia NOC from scratch

Cut MTTR (mean time to repair - how long it takes to fix an incident) by 20%. Built a failure-forecast model on historical data and prevented 4+ major network incidents a year.

2021-2023 · Bank

Led Problem Management and monitoring in a digital transformation

Owned Problem Management (the root-cause practice) plus 12 more ITIL practices. Built a monitoring service from zero. Held the SLA at 99.9% ± 0.2%.

2019-2021 · ITSM consulting

Independent infrastructure audits

Cut the failure rate by roughly 50% across engagements. Raised availability from 95% to 99.8%.

1999-2004 · Lensvyaz

Chief engineer - telecom, networks, power

Radio communications engineer by training. "Master of Communications" award, 2013.

The engine looks for statistical patterns. I propose the fix.

The software and AI system, built on a branching analysis algorithm, shows what happened in your data and what may happen next. My task is to interpret the results correctly, drawing on my experience and the experience of my colleagues, and to propose the best plan of action - one based on current practice and not requiring excessive resources.

What our statistical module looks for

Eight core methods out of 20+ in the OpsLab engine. The full list also covers correlations, trend diagnostics, anomaly detection and a per-service breakdown.

Priority
Priority (Pareto ranking)
Which 20% of incidents cause 80% of the downtime, so you fix the few that matter. The Gini coefficient (a measure of inequality: 0 - all incidents equally bad, 1 - one incident caused everything) shows how concentrated the problem is.
3 of 90 incidents = 41% of all downtime. Gini = 0.67.
Root cause
Root causes (clustering)
Groups incidents by behaviour - time of day, duration, frequency - instead of by ticket category. A long list of "different" errors often turns out to be a small number of root causes.
90 incidents -> 2 clusters: "night, short" and "morning, long". Two causes, two teams.
Cycles
Cycles (periodicity)
Finds hidden repeat cycles in uneven time series - a method from astrophysics applied to incident logs. It also shows in which hours and days failures cluster. It reports FAP (false alarm probability - the chance the cycle is random noise).
Period = 23.1 h, FAP = 4x10⁻⁸ -> almost certainly a scheduler, not chance.
Verification
Before / after a change
Tests whether a deploy or a config change really moved the incident rate, and by how much. Cliff's delta (effect size - how big the shift is, from 0 to 1) turns "it feels worse" into a number.
Cliff's delta = +0.445, p < 0.001 -> the rate grew, confidence above 99.9%.
Forecast
Forecast and tail risk
Estimates how many failures to expect in the coming weeks, and the probability of a rare but very long outage - the one that breaks an SLA. The answer is always a range, never a falsely precise single number.
7 days - about 67 incidents (52-84). Chance of an outage longer than 6 hours - 7%.
Event links
Event links
Finds which events appear together and in what order, so an early alert can sit on the leading signal instead of the crash itself. Lift shows how many times more often a pair occurs than by chance.
database_timeout -> app_crash in 89% of cases, lift = 4.2.
Reliability
MTBF and MTTR without an agent
Mean time between failures (MTBF) and mean time to repair (MTTR) straight from your export - no agent, no access to your servers. Mean, median and P95 (the value only 5% of cases are worse than) together show the bad day, not only the typical one.
MTBF: 2.0 h mean, 0 h median, 24 h P95 - rare long pauses pull the mean up.
Benchmark
Industry norms
Puts your metrics next to the thresholds of your industry. The thresholds are data, not a hard-coded list. Where a threshold is unknown, the engine says so instead of inventing a norm.
P95 response 340 ms against a 200 ms telecom threshold; jitter 12 ms - within the 30 ms norm.
What comes next
Popular fixes after the analysis
Which problems we find most often and what to do about them, in ITIL / ITSM terms.

How we work together

Three ways to start. Try something small first - no obligation to continue.

1

Diagnostics -> retainer -> project

We start by reviewing your incident history: what causes most downtime, what repeats, what to fix first. If it proves useful, we continue on a regular basis or as a separate project - for example, setting up Problem Management or a NOC.

2

A run inside your perimeter

If your data cannot leave the company, the analysis runs on your own machine. You keep the data - we only get the report, and we go through it together.

3

Interim / part-time role

We take on the reliability role for a few months: a fixed share of the week, an agreed scope, a clear end date. Useful while you are hiring a permanent specialist or going through a period of change.

Let's discuss your incidents

Just name the main problem you are facing, and we will take it from there.