Skip to content
OperationsTelia

Telia's first end-to-end monitoring and observability platform

Establishing one monitoring and observability platform across BSS, OSS, applications, infrastructure, databases and network — using Splunk, Nagios and BMC TrueSight — so operations could see services rather than components.

  • Splunk
  • Nagios
  • BMC TrueSight
  • Event management
  • Observability
  • ITSM

Context

What was happening?

Monitoring at Telia had accumulated by domain: infrastructure tooling for servers and network, application-specific checks, ITSM for process. Each gave a partial view, and the picture of a customer-facing service was assembled by people during an incident rather than held by a system.

Challenge

What problem needed solving?

Establish a single end-to-end monitoring and observability capability spanning BSS, OSS, applications, infrastructure, databases and network — the first of its kind at Telia — without replacing every domain tool or stalling live operations.

  • Events and metrics from six technology domains with different owners and formats.
  • No shared model connecting an infrastructure alarm to the BSS/OSS service it affected.
  • Existing tooling that had to be integrated, not discarded.

Architecture

What was the architectural approach?

The platform combined Splunk, Nagios and BMC TrueSight — analytics across logs, metrics and events; availability monitoring; and event and infrastructure management — into one capability spanning all six domains, so that a cross-domain picture existed above the individual tools rather than inside any one of them.

Monitoring and observability layers
  1. 01Domain sources

    BSS, OSS, applications, infrastructure, databases, network

  2. 02Monitoring & event management

    Nagios and BMC TrueSight

  3. 03Observability & analytics

    Splunk — logs, metrics and events across domains

  4. 04Service view

    Impact expressed against services, feeding ITSM

My role

What I actually did

Established the platform: defining the monitoring architecture across the six domains, the role of each tool, the integration between them and the operating model for keeping coverage and the service model current as the estate changed.

Transformation

What changed?

Operations moved from domain-by-domain alarm handling to a single cross-domain view. Ownership of monitoring coverage and the service model became an explicit responsibility rather than tribal knowledge — the part of an assurance change most often left out.

Outcome

What was achieved?

Telia's first end-to-end monitoring and observability platform, spanning BSS, OSS, applications, infrastructure, databases and network. Quantified operational results are not published here.

What I learned

The broader architectural insight

  • Give each tool one role; overlapping tools produce overlapping truths.
  • Correlation without a service model is better-organised noise.
  • ITSM and monitoring should share a model, not just a ticket integration.
  • Service assurance in cloud-native networks

    When network functions are software on shared infrastructure, the mapping from resource to service becomes dynamic. Assurance has to follow the model, not the box.

  • What Agentic NOC means for OSS

    An agent that correlates events and proposes remediation depends entirely on the service model, inventory and telemetry the OSS gives it. Agentic NOC is an OSS programme before it is an AI programme.