Telia's first end-to-end monitoring and observability platform
Establishing one monitoring and observability platform across BSS, OSS, applications, infrastructure, databases and network — using Splunk, Nagios and BMC TrueSight — so operations could see services rather than components.
- Splunk
- Nagios
- BMC TrueSight
- Event management
- Observability
- ITSM
Context
What was happening?
Monitoring at Telia had accumulated by domain: infrastructure tooling for servers and network, application-specific checks, ITSM for process. Each gave a partial view, and the picture of a customer-facing service was assembled by people during an incident rather than held by a system.
Challenge
What problem needed solving?
Establish a single end-to-end monitoring and observability capability spanning BSS, OSS, applications, infrastructure, databases and network — the first of its kind at Telia — without replacing every domain tool or stalling live operations.
- Events and metrics from six technology domains with different owners and formats.
- No shared model connecting an infrastructure alarm to the BSS/OSS service it affected.
- Existing tooling that had to be integrated, not discarded.
Architecture
What was the architectural approach?
The platform combined Splunk, Nagios and BMC TrueSight — analytics across logs, metrics and events; availability monitoring; and event and infrastructure management — into one capability spanning all six domains, so that a cross-domain picture existed above the individual tools rather than inside any one of them.
- 01Domain sources
BSS, OSS, applications, infrastructure, databases, network
- 02Monitoring & event management
Nagios and BMC TrueSight
- 03Observability & analytics
Splunk — logs, metrics and events across domains
- 04Service view
Impact expressed against services, feeding ITSM
My role
What I actually did
Established the platform: defining the monitoring architecture across the six domains, the role of each tool, the integration between them and the operating model for keeping coverage and the service model current as the estate changed.
Transformation
What changed?
Operations moved from domain-by-domain alarm handling to a single cross-domain view. Ownership of monitoring coverage and the service model became an explicit responsibility rather than tribal knowledge — the part of an assurance change most often left out.
Outcome
What was achieved?
Telia's first end-to-end monitoring and observability platform, spanning BSS, OSS, applications, infrastructure, databases and network. Quantified operational results are not published here.
What I learned
The broader architectural insight
- Give each tool one role; overlapping tools produce overlapping truths.
- Correlation without a service model is better-organised noise.
- ITSM and monitoring should share a model, not just a ticket integration.
Related thinking
Service assurance in cloud-native networks
When network functions are software on shared infrastructure, the mapping from resource to service becomes dynamic. Assurance has to follow the model, not the box.
What Agentic NOC means for OSS
An agent that correlates events and proposes remediation depends entirely on the service model, inventory and telemetry the OSS gives it. Agentic NOC is an OSS programme before it is an AI programme.