WAN Design Requirements

DMVPN

NETWORK INSIGHT · ENGINEERING GUIDE

DMVPN: Designing and Investigating Dynamic VPN Overlays

Dynamic Multipoint VPN (DMVPN) combines multipoint GRE, Next Hop Resolution Protocol (NHRP), IPsec and dynamic routing to build scalable VPN connectivity across an underlying IP network.

The engineering challenge is not simply creating a secure tunnel. A DMVPN deployment must dynamically discover remote endpoints, establish the appropriate tunnel path, maintain routing information and provide efficient communication between sites as the network grows.

This guide moves beyond configuration syntax and examines DMVPN as an operational system. You will follow the relationship between the underlay and overlay, observe how NHRP provides endpoint resolution, investigate how spoke-to-spoke connectivity develops, and correlate routing, tunnel and security evidence when troubleshooting the network.

01 · SEE DMVPN Architecture Understand the underlay, overlay, hub, spokes and core DMVPN components.
02 · OBSERVE NHRP & Tunnel Evidence Examine mappings, tunnel state, routing information and security evidence.
03 · INVESTIGATE Spoke-to-Spoke Paths Follow NHRP redirect and resolution as traffic moves toward a direct path.
04 · OPERATE Routing & Resiliency Correlate routing, IPsec and failure evidence to verify network health.
Investigation Focus
Follow the evidence

The interactive sections turn DMVPN behaviour into an engineering investigation: identify the path, inspect the evidence, determine what changed and verify the resulting network state.

Engineering Objective

By the end of the guide, you should be able to explain how DMVPN forms and optimises its overlay, identify the evidence created by NHRP and IPsec, distinguish hub-and-spoke from direct spoke-to-spoke forwarding, and use multiple observations to isolate a DMVPN failure.

The interactive sections are designed as engineering exercises rather than demonstrations. Each stage builds on the previous one so that architecture, evidence, investigation and operational verification form one continuous learning path.
ENGINEERING GUIDE · 01 — SEE

Understand the DMVPN Architecture

Start with the dependency chain. DMVPN is not a single protocol: the underlay, mGRE, NHRP, IPsec and routing functions each solve a different engineering problem. Before troubleshooting traffic, establish which layer is responsible for connectivity, discovery, security, routing and forwarding.

UNDERLAY mGRE NHRP IPsec ROUTING
ENGINEERING EVIDENCE · SECTION 01

SEE — DMVPN Architecture

Before investigating a DMVPN failure, understand how the underlay, mGRE, NHRP, IPsec and routing functions fit together to create the overlay.

DMVPN Architecture Stack
UNDERLAY
IP connectivity between the physical / NBMA endpoints
REACHABILITY
mGRE
Multipoint tunnel framework for dynamic destinations
OVERLAY
NHRP
Dynamic discovery and logical-to-NBMA mapping
CONTROL
IPsec
Encryption and protection of overlay traffic
SECURITY
ROUTING
Destination reachability and path selection
CONTROL
Underlay

The underlying IP network provides basic reachability between DMVPN routers. If the underlay cannot reach the required peer, the overlay cannot form correctly.

From Infrastructure To Application
01
UNDERLAY
IP / NBMA reachability
›
02
mGRE
Multipoint overlay
›
03
NHRP
Peer discovery
›
04
IPsec
Traffic protection
›
05
ROUTING
Path selection
›
06
APPLICATION
End-to-end traffic
DMVPN Traffic Models
PHASE 1

Hub-and-Spoke

Traffic remains logically centred on the hub. The hub provides the central connectivity point for the spokes.

PHASE 2

Spoke-to-Spoke

Spokes can dynamically establish direct connectivity when the routing and NHRP information support the destination.

PHASE 3

Hub-Initiated

The hub can participate in redirecting traffic toward a more efficient spoke-to-spoke path.

Engineering View

DMVPN should not be investigated as a single protocol. A tunnel can appear operational while NHRP mappings, IPsec security associations or routing information are incorrect.

The investigation therefore follows the dependency chain: Underlay → mGRE → NHRP → IPsec → Routing → Forwarding.

ENGINEERING QUESTIONS
Can the underlay reach the peer?
Does NHRP know where the peer is?
Are IPsec security associations present?
Is routing selecting the expected path?
Does forwarding match the design?
ENGINEERING EVIDENCE · DMVPN ARCHITECTURE

SEE — DMVPN Architecture

Interrogate the DMVPN dependency chain. Select a layer to see what it does, what evidence proves it is working, and what depends on it.
ARCHITECTURE ONLINE
DMVPN Dependency Stack
01
UNDERLAY
Physical/IP transport and NBMA reachability
REACHABLE
02
mGRE
Multipoint tunnel framework
UP
03
NHRP
Logical tunnel-to-NBMA peer mapping
MAPPED
04
IPsec
Protection of tunnel traffic
SA UP
05
ROUTING
Prefix reachability and path selection
VALID
Engineering Dependency Chain
UNDERLAY › mGRE › NHRP › IPsec › ROUTING
LAYER 01 · TRANSPORT

Underlay Reachability

The underlay provides IP reachability between the physical or NBMA endpoints. DMVPN cannot establish useful overlay behaviour if the transport network cannot reach the required peers.

Function Transport
Example 198.51.100.11
Evidence IP route / ping
Failure boundary Peer unreachable
DMVPN Operating Model
Phase 1 · Hub-and-Spoke Spokes primarily communicate through the hub.
Phase 2 · Spoke-to-Spoke Spokes can establish direct dynamic paths.
Phase 3 · Hub-Assisted Hub redirects traffic towards an optimized peer path.
Evidence Command
show ip route 198.51.100.1
Selected Layer Depends On
The underlay is the foundation. Higher DMVPN functions ultimately depend on transport reachability.
IP Reachability NBMA Transport
What This Layer Does Not Prove
A reachable underlay does not prove that the DMVPN tunnel, NHRP mappings, IPsec SAs or routing are correct.
mGRE NHRP IPsec Routing
ENGINEERING GUIDE · 02 — DISCOVER

How NHRP Builds Dynamic Peer Knowledge

Once the tunnel framework exists, DMVPN needs a way to associate logical tunnel addresses with real NBMA destinations. NHRP provides that discovery and mapping function, allowing spokes and the NHS to build the peer information required by the overlay without turning NHRP itself into the routing protocol.

REGISTRATION NHS / NHC MAPPING RESOLUTION
ENGINEERING EVIDENCE · SECTION 02

DISCOVER — NHRP and Dynamic Peer Mapping

mGRE provides the multipoint tunnel framework, but it does not tell a DMVPN router where another peer exists on the underlay. NHRP provides the discovery and mapping mechanism that connects the logical tunnel address with the peer's NBMA address.

NHRP Discovery Sequence
01
Spoke Registration
The spoke tells the NHS where it can be reached.
02
NHRP Mapping
The hub builds dynamic peer information.
03
Resolution Request
The destination peer's location is requested.
04
Resolution Reply
Logical and NBMA information is returned.
Spoke Registration
REGISTER
SPOKE 1
NHC
HUB
NHS
SPOKE 2
NHC
NHRP EVENT
The spoke registers its tunnel address and NBMA address with the NHS.
Dynamic NHRP Information

NHRP Mapping View

DEVICE ROLE TUNNEL NBMA STATE
HUB NHS 172.16.100.1 198.51.100.1 LOCAL
SPOKE 1 NHC 172.16.100.11 198.51.100.11 REGISTERED
SPOKE 2 NHC 172.16.100.12 198.51.100.12 PENDING

Engineering Interpretation

NHRP is not the routing protocol that determines which destination network should be used. Its role is to provide the information required to reach another DMVPN peer by relating a logical tunnel address to an NBMA address.

This makes NHRP an important dependency when investigating dynamically established DMVPN connectivity.

CLI EVIDENCE
show ip nhrp
show ip nhrp nhs
show dmvpn
show ip interface tunnel 0
ENGINEERING CHECKPOINT
Is the spoke registered with the correct NHS?
Does the NHRP table contain the expected NBMA mapping?
Can the discovered peer be reached across the underlay?
CORE IDEA
NHRP discovers where a DMVPN peer lives on the underlay.
Routing still determines destination reachability and path selection.
ENGINEERING EVIDENCE · NHRP DISCOVERY

DISCOVER — Dynamic Peer Mapping

Step through the NHRP control-plane sequence and watch how a DMVPN peer moves from registration to a usable logical-to-NBMA mapping.
NHRP CONTROL PLANE
Live NHRP Discovery Sequence
The hub acts as the NHS. Spokes register their tunnel/NBMA information and can subsequently resolve dynamic peer information.
SPOKE 1
NHC
172.16.100.11
198.51.100.11
HUB
NHS
172.16.100.1
198.51.100.1
SPOKE 2
NHC
172.16.100.12
198.51.100.12
INITIAL NHRP STATE
EVENT 01 · REGISTRATION
Spoke 1 sends NHRP registration information towards the NHS, advertising its logical tunnel address and NBMA address.
Control-Plane Events
01
Spoke Registration NHC registers its logical and NBMA information with the NHS.
02
NHRP Mapping The NHS records the logical-to-NBMA relationship.
03
Resolution Request A spoke requests information for a dynamic peer.
04
Resolution Reply The requested peer mapping becomes available.
NHRP Mapping Table
Node Logical NBMA State
HUB 172.16.100.1 198.51.100.1 LOCAL
SP1 172.16.100.11 198.51.100.11 REGISTERED
SP2 172.16.100.12 198.51.100.12 PENDING
Engineering Interpretation

NHRP is not a routing protocol. Its role here is to provide the information needed to map a logical DMVPN peer to its NBMA transport address.

Operational Evidence
show ip nhrp show ip nhrp nhs show dmvpn show ip interface tunnel 0
What The Mapping Tells You
Logical address → NBMA address
The mapping establishes how a DMVPN logical peer relates to its underlying transport endpoint. It does not, by itself, prove that IPsec is established or that the routing table contains the correct destination path.
ENGINEERING GUIDE · 03 — PROTECT

How IPsec Protects the DMVPN Overlay

NHRP can identify a peer, but discovery alone does not provide security. IPsec establishes the protection required for DMVPN traffic, securing the encapsulated overlay between peers. A tunnel interface being operational is therefore not sufficient evidence that encrypted traffic is successfully passing.

IKE AUTHENTICATION IPsec SA ESP ENCRYPTION
ENGINEERING EVIDENCE · SECTION 03

PROTECT — IPsec Security

NHRP can identify where a DMVPN peer exists, but discovery does not provide confidentiality or integrity. IPsec supplies the security layer that protects DMVPN traffic as it crosses the underlying network.

Security Establishment Sequence
01
IKE / Negotiation
Security parameters are negotiated between the peers.
02
Authentication
Each peer verifies the identity of the other side.
03
IPsec SA
The negotiated security association is established.
04
GRE Protection
Overlay traffic is protected before entering the underlay.
05
Protected Transport
The secured packet traverses the underlying IP network.
IKE / Security Negotiation
NEGOTIATE
DMVPN PEER A
IPsec PEER
DMVPN PEER B
IPsec PEER
UNDERLAY / NBMA NETWORK
IPSEC PROTECTED
SECURITY EVENT
The peers negotiate the parameters required to establish protected communication.
What The Router Should Show

Security Association Evidence

PEER
198.51.100.12
STATE
ESTABLISHED
PROTOCOL
ESP

Engineering Interpretation

A working tunnel interface does not by itself prove that protected traffic can pass. IPsec security associations must be established and traffic must be processed by the correct security policy.

When troubleshooting, separate the questions: Is the peer reachable? Is the security association established? Are packets actually being encrypted and decrypted?

CLI EVIDENCE
show crypto isakmp sa
show crypto ipsec sa
show crypto session
ENGINEERING EVIDENCE · IPSEC SECURITY

PROTECT — IPsec Security

Follow the security establishment process from IKE negotiation through an operational IPsec SA and protected DMVPN traffic.

SECURITY STANDBY
Live Security Path
PEER A ↔ PEER B
STAGE 1 · IKE
A
DMVPN PEER A
198.51.100.11
B
DMVPN PEER B
198.51.100.12
UNDERLAY / NBMA TRANSPORT
NBMA · 198.51.100.11 → 198.51.100.12
IKE negotiation begins The peers establish the security negotiation required before protected traffic can be exchanged.
Security Establishment
CLICK TO INSPECT
SA Telemetry
LIVE STATE
Peer
198.51.100.12
State
NEGOTIATING
Protocol
IKE
Encrypt
0 pkts
Decrypt
0 pkts
Protection
NOT ACTIVE
Selected Stage Evidence

IKE establishes the negotiation context. At this point, an operational IPsec SA has not yet been demonstrated.

R1# show crypto isakmp sa
Engineering Interpretation

A DMVPN tunnel interface being operational does not by itself prove that encrypted traffic is successfully passing between peers.

ENGINEERING PRINCIPLE: NHRP identifies the peer's NBMA information; IPsec protects the traffic. When troubleshooting, separate peer discovery, security-association establishment and actual encrypted packet counters rather than treating “tunnel up” as proof of end-to-end operation.
ENGINEERING GUIDE · 04 — ROUTE

How Routing Determines the Overlay Path

DMVPN provides the overlay connectivity, but routing protocols determine which destination prefixes are reachable and which path is preferred. EIGRP, OSPF and BGP can operate across the overlay, each applying its own path-selection logic.

EIGRP OSPF BGP BEST PATH CEF
ENGINEERING EVIDENCE · SECTION 04

ROUTE — Routing Across the DMVPN Overlay

DMVPN provides the overlay connectivity, but routing determines which destination prefixes are reachable and which path the router prefers. EIGRP, OSPF and BGP can operate across the overlay, with the routing design influencing how traffic moves between spokes and hubs.

Select A Routing Model
EIGRP
Fast convergence and native Cisco integration.
IGP
OSPF
Link-state routing across the overlay.
IGP
BGP
Policy-driven routing and path control.
EGP
Route Decision View

EIGRP Routing Model

EIGRP can exchange routes across the DMVPN overlay and is commonly used in Cisco-centric DMVPN designs. Its metrics and next-hop behaviour influence the selected path.

Metric
LOWER
Feasibility
CHECK
Next Hop
OVERLAY
Split Horizon
DESIGN
Candidate Paths
ROUTE INSTALLED
PATH A · DIRECT SPOKE
Via Tunnel0 · metric 90
→
10.40.40.0/24
Preferred path
PATH B · HUB
Via Tunnel0 · metric 120
→
10.40.40.0/24
Alternate path
PATH C · BACKUP
Via secondary path
→
10.40.40.0/24
Standby
Select a candidate path to inspect the routing decision.

Engineering Interpretation

A DMVPN tunnel being operational does not mean that the expected route will be installed. The overlay provides connectivity between routers; the routing protocol decides which prefixes are reachable and which path should be preferred.

Investigation Question

If the destination prefix exists but traffic takes the wrong path, investigate the routing information before assuming that the DMVPN tunnel itself is broken.

CLI EVIDENCE
show ip route
show ip route 10.40.40.0
show ip protocols
show ip cef 10.40.40.0

Following the Packet Through the Overlay

A DMVPN tunnel is not the packet itself. It is the transport framework that allows one router to carry an original IP packet across an underlying network. To understand forwarding behaviour, separate the inner packet from the DMVPN overlay and the underlay transport.

Consider a packet travelling from a host behind Spoke 1 to a host behind Spoke 2. The original packet might have the source 10.10.10.10 and destination 10.40.40.40. Those addresses describe the actual communication. They are not replaced simply because the packet enters the DMVPN tunnel.

Source Host → mGRE Tunnel → IPsec Protection → Underlay → Remote Peer → Destination

The Inner Packet

The original IP packet contains the addresses used by the overlay routing decision and, ultimately, the destination host. For example:

# Original IP packet
SRC       10.10.10.10
DST       10.40.40.40
PROTOCOL  IP

Spoke 1 uses its routing and forwarding information to determine that the destination prefix is reachable through the DMVPN overlay. The packet is therefore handed towards the tunnel interface rather than being forwarded directly using the physical underlay interface.

mGRE Creates the Overlay Transport

mGRE provides the multipoint tunnel framework. Instead of requiring a separate permanent tunnel interface for every possible peer, a DMVPN router can use a single multipoint GRE tunnel to support dynamically discovered peers.

The important engineering distinction is that mGRE provides the tunnel mechanism, while NHRP provides the peer discovery and logical-to-NBMA mapping information required to reach dynamic peers.

Inner

Original source and destination addresses remain associated with the actual IP communication.

Overlay

mGRE carries the original packet across the DMVPN tunnel infrastructure.

Underlay

The transport network carries the encapsulated traffic between the DMVPN peers.

IPsec Protects the Tunnel Traffic

After the tunnel traffic has been constructed, IPsec provides the security layer. Depending on the deployment, the GRE traffic is protected using IPsec security associations and ESP.

This creates an important troubleshooting boundary. A tunnel interface can appear operational while the required IPsec security association is missing or while protected traffic is not being successfully encrypted and decrypted.

Engineering Evidence

Tunnel up does not automatically mean traffic is working. Validate the tunnel, NHRP state, IPsec security associations, routing information and forwarding behaviour independently.

The Underlay Sees the Outer Transport

Once the packet has been encapsulated and protected, the underlay network forwards the resulting transport traffic. The underlay is concerned with reaching the remote DMVPN peer rather than understanding the final application destination inside the overlay.

# Example outer transport addresses
OUTER SRC  198.51.100.11
OUTER DST  198.51.100.12
PROTOCOL   IPsec / ESP

This distinction is extremely useful during troubleshooting. If the underlay cannot reach 198.51.100.12, the overlay cannot carry the packet regardless of whether the routing configuration itself is correct.

Decapsulation at the Remote Peer

At the remote DMVPN router, the process is reversed. The protected transport is received and processed by IPsec. The tunnel encapsulation is then removed, exposing the original IP packet.

The remote router can then perform normal IP forwarding towards 10.40.40.40. The final destination therefore sees the original IP communication rather than the intermediate NBMA transport addresses used by the DMVPN infrastructure.

What Each Layer Actually Knows

Overlay Routing

Determines how the destination prefix is reached through the DMVPN overlay.

IPsec

Protects the tunnel traffic and maintains the required security associations.

Underlay

Provides transport reachability between the physical/NBMA endpoints.

Engineering Principle

Follow the packet, not just the tunnel. When a DMVPN path fails, determine which layer can no longer carry the packet. A working underlay does not prove NHRP is correct. An established NHRP mapping does not prove IPsec is protecting traffic. An established IPsec SA does not prove the routing table contains the expected destination path.

Evidence to Collect

The data plane can be correlated with the control-plane state using a small set of operational commands:

# Forwarding decision
show ip cef 10.40.40.40

# Tunnel state and counters
show interfaces tunnel 0

# NHRP peer mappings
show ip nhrp

# IPsec protection and packet counters
show crypto ipsec sa

These outputs should be interpreted together. The goal is not simply to prove that individual commands return output, but to establish a continuous forwarding chain from the source host to the destination.

ENGINEERING EVIDENCE · OVERLAY ROUTING

ROUTE — Routing Across the DMVPN Overlay

Compare candidate paths and see how the routing protocol determines which overlay path becomes the forwarding decision.

ROUTE ANALYSIS READY
Candidate Route Topology
10.40.40.0/24
R1
SOURCE
10.10.10.0/24
RR
RR1
AS65001
R4
DESTINATION
10.40.40.0/24
Candidate paths discovered The same destination prefix is visible through multiple overlay paths. The routing process must select the usable best path.
Routing Protocol
SELECT MODEL
Candidate Paths
CLICK TO INSPECT
Decision Telemetry
LIVE ANALYSIS
Protocol
EIGRP
Selected
PATH A
Destination
10.40.40.0
FIB State
CANDIDATES
Routing Evidence

The routing table contains multiple candidate paths to the same destination. The selected path is the route that will be installed for forwarding.

R1# show ip route 10.40.40.0
R1# show ip cef 10.40.40.0
Engineering Interpretation

DMVPN provides the overlay connectivity. The routing protocol determines which destination prefix is reachable and which path should be preferred.

ENGINEERING PRINCIPLE: A healthy DMVPN tunnel does not automatically mean the correct route is being used. Separate overlay connectivity from route selection, then verify the installed RIB and CEF forwarding decision.
ENGINEERING GUIDE · 05 — FORWARD

FORWARD — Follow the DMVPN Data Plane

Once the overlay is built, the real engineering question is simple: how does an actual packet travel? Follow the original IP packet from the source host through the DMVPN tunnel, across the underlay, and back into the remote network. This separates the inner packet, overlay tunnel, IPsec protection, and underlay transport so failures can be isolated to the correct layer.

INNER PACKET mGRE IPsec UNDERLAY DECAPSULATION
ENGINEERING EVIDENCE · SECTION 05

FORWARD — DMVPN Data Plane & Packet Flow

Once the overlay is established and routing has selected a path, the next question is simple: what actually happens to the packet? The DMVPN data plane combines the original IP packet, the tunnel overlay, IPsec protection and the underlay transport. Understanding these layers makes it possible to identify exactly where forwarding succeeds or fails.

→
Engineering principle The original packet is not replaced by the DMVPN transport. It is encapsulated, protected and carried across the underlay, then decapsulated at the remote endpoint.
01
Source Host Original IP packet is generated.
02
Spoke Routing selects Tunnel0.
03
mGRE Overlay encapsulation is added.
04
IPsec Tunnel traffic is protected.
05
Underlay NBMA transport forwards it.
06
Remote Spoke Traffic is decapsulated.
Packet Construction

One packet, multiple forwarding views

The endpoint application creates an ordinary IP packet. The DMVPN tunnel then provides the transport framework around that packet. mGRE identifies the tunnel overlay while IPsec protects the resulting tunnel traffic before it crosses the physical network.

INNER 10.10.10.10 → 10.40.40.40
TUNNEL Spoke1 Tunnel0 → Spoke2 Tunnel0
OUTER 198.51.100.11 → 198.51.100.12
Forwarding Layers

What each layer actually sees

IP
Inner IP The original source and destination prefixes used by the overlay routing decision.
GRE
mGRE Overlay Provides the tunnel framework used to carry the original packet between DMVPN peers.
SEC
IPsec Protection Protects the tunnel traffic before it enters the physical underlay.
WAN
Underlay Forwards the transport using the outer NBMA addresses.
Remote Processing

Decapsulation restores the original forwarding context

At the remote DMVPN peer, the protected transport is processed in the reverse direction. IPsec protection is removed, the mGRE encapsulation is removed, and the original IP packet becomes available for normal forwarding toward its destination. This distinction is critical: the physical network forwards the outer transport, while the DMVPN overlay forwards the inner destination prefix.

CEF Decision

Confirm which interface and next-hop the router selected for the destination prefix.

show ip cef 10.40.40.0

Tunnel State

Confirm the tunnel interface is operational and carrying the expected traffic.

show interfaces tunnel 0

IPsec Counters

Confirm encryption and decryption counters are increasing when traffic is generated.

show crypto ipsec sa
Engineering interpretation

If the underlay is reachable but the destination cannot be forwarded, do not treat the problem as a single “DMVPN tunnel failure.” Separate the investigation into the layers: underlay transport → tunnel encapsulation → IPsec protection → overlay routing → forwarding. The point at which the packet stops progressing identifies the most useful next piece of evidence.

ENGINEERING EVIDENCE · SECTION 05

FORWARD — Follow the DMVPN Data Plane

Trace one packet from the source host to the remote destination. Select each stage to see which address, header and forwarding function is active at that point in the journey.
DATA PLANE READY
Live Packet Journey
PACKET
01
SOURCE Original IP packet
02
mGRE Overlay encapsulation
03
IPsec Protected transport
04
UNDERLAY Outer NBMA forwarding
05
DESTINATION Decapsulation + delivery
Source Host
10.10.10.10
→
Remote Host
10.40.40.40

Original IP Packet

The source creates the original IP packet. At this point there is no DMVPN tunnel header and the destination is the final host.

Stage SOURCE
Source 10.10.10.10
Destination 10.40.40.40
Forwarding HOST
Engineering Evidence
show ip cef 10.40.40.40 show interfaces tunnel 0 show crypto ipsec sa
Packet Header Inspection INNER
Inner IP
SRC 10.10.10.10
DST 10.40.40.40
DMVPN Overlay
Tunnel0
mGRE / NHRP peer
Outer Transport
NBMA 198.51.100.11
→ 198.51.100.12
STEP 01
Original Packet Host creates IP traffic.
STEP 02
mGRE Overlay encapsulation begins.
STEP 03
IPsec GRE traffic is protected.
STEP 04
Underlay Outer addresses are forwarded.
STEP 05
Decapsulation Remote peer removes tunnel layers.
RESULT
Forwarded Original packet reaches destination.
NETWORK INSIGHT · INTERACTIVE ENGINE

DMVPN — Phase 3 Explorer

Explore DMVPN Phase 3, follow the NHRP Redirect and Resolution process, and see how the resulting spoke-to-spoke path is used.
Phase 3 Network Topology
PHASE 3 · NHRP REDIRECT
PHASE 3 · NHRP REDIRECT
Spoke 1
DMVPN SPOKE
Hub
NHRP NHS
Spoke 2
DMVPN SPOKE
NHRP REDIRECT
STEP 1 / 4
Initial traffic crosses the Hub
Spoke 1 initially sends traffic toward Spoke 2 through the DMVPN hub.
Phase 3 Detail
NHRP · OPTIMIZATION
Phase 3 — NHRP Redirect
The hub can signal that traffic between two spokes can use a more direct path.
Initial Path Spoke 1 → Hub → Spoke 2
NHRP Redirect + Resolution
Direct Path Yes
NHRP Control Plane
REGISTRATION · RESOLUTION
NHRP · CONTROL PLANE
HUB
NHRP NHS
SPOKE 1
NHRP CLIENT
SPOKE 2
NHRP CLIENT
REGISTRATION
REGISTRATION
REMOTE MAPPING / RESOLUTION
Selected Object
NHRP
NHRP Control Plane
NHRP allows DMVPN spokes to register their tunnel information and discover remote spoke mappings.
Function Peer discovery
Server Hub / NHS
Security IPsec
DMVPN Packet Encapsulation
GRE · IPSEC · UNDERLAY
ORIGINAL IP PACKET
Spoke 1
SOURCE
Spoke 2
DESTINATION
Original IP Packet
Source → Destination
GRE / mGRE
Overlay tunnel encapsulation
IPsec
Authentication / encryption protection
Underlay / Transport
Physical or routed IP transport
The original packet is carried through the DMVPN overlay and protected by IPsec.
Encapsulation Detail
PACKET FLOW
Original IP Packet
The original application packet is generated by the source network before tunnel encapsulation.
Layer Payload
Technology IP
Purpose Original traffic
ENGINEERING GUIDE · 06 — INVESTIGATE

INVESTIGATE — DMVPN Failure Workflow

A DMVPN failure should be investigated as a dependency chain, not as a single tunnel problem. Work from the underlay through mGRE, NHRP, IPsec, routing and finally forwarding to isolate the first broken layer.

Underlay mGRE NHRP IPsec Routing Forwarding
ENGINEERING EVIDENCE · SECTION 06

INVESTIGATE — DMVPN Failure Workflow

Troubleshooting DMVPN becomes much easier when the investigation follows the dependency chain established throughout this guide. Instead of treating “the tunnel is down” as the diagnosis, identify the first layer that has failed and prove it with operational evidence.

!
The first failed dependency is usually more valuable than the final symptom.

A missing route may be caused by routing policy, an unresolved NHRP peer, an unavailable IPsec security association, or a broken underlay. The correct troubleshooting sequence prevents a downstream symptom from being mistaken for the root cause.

Start With the Dependency Chain

DMVPN is not a single protocol. It is a collection of functions operating together. Each layer depends on the layers beneath it, so investigation should normally move from basic transport reachability toward the final forwarding decision.

01

Underlay

Confirm that the NBMA transport addresses are reachable before investigating the overlay.

show ip route
02

mGRE

Verify that Tunnel0 is operational and that the tunnel interface has the expected state and addressing.

show interfaces tunnel 0
03

NHRP

Check whether logical tunnel peers are correctly registered and mapped to their NBMA addresses.

show ip nhrp
04

IPsec

Confirm that the required security associations exist and that protected traffic is being encrypted and decrypted.

show crypto ipsec sa
05

Routing

Determine whether the destination prefix exists and whether the expected path has been selected.

show ip route
06

Forwarding

Verify that the selected route produces the expected CEF forwarding decision and next hop.

show ip cef

Common Failure Signatures

The same user-visible symptom can originate from different layers. The objective is therefore to correlate the symptom with the control-plane and forwarding evidence underneath it.

Underlay Unreachable

The NBMA address of the remote peer cannot be reached. NHRP, IPsec and overlay routing may subsequently appear broken because their transport dependency is unavailable.

NHRP Mapping Missing

The tunnel framework may exist, but the required logical-to-NBMA peer information is absent or incomplete. Dynamic peer communication can therefore fail.

IPsec SA Missing

Peer discovery may succeed while protected traffic still fails because the required security association has not been established or is not passing traffic.

Routing Mismatch

The DMVPN overlay can be operational while the routing process selects an unexpected path, rejects a prefix, or fails to install the destination in the forwarding table.

Evidence Before Diagnosis

A strong troubleshooting process does not begin by changing configuration. First establish what the network believes is happening. Then compare that evidence with the expected dependency state.

Peer and Tunnel Evidence

Establish whether the transport, tunnel and NHRP control plane agree about the remote peer.

! Underlay
show ip route 198.51.100.12
ping 198.51.100.12

! DMVPN / NHRP
show dmvpn
show ip nhrp

Security and Routing Evidence

Once peer reachability is established, determine whether protected transport and route installation are functioning.

! IPsec
show crypto ipsec sa
show crypto session

! Routing / forwarding
show ip route 10.40.40.0
show ip cef 10.40.40.0

The Correct Troubleshooting Order

The investigation should move from the dependency that everything else requires toward the final forwarding decision. Skipping directly to routing or application symptoms can produce misleading conclusions.

UNDERLAY
›
mGRE
›
NHRP
›
IPsec
›
ROUTING
›
FORWARDING

Engineering Conclusion

A DMVPN investigation should move from reachability → discovery → security → routing → forwarding. Each layer answers a different question. Underlay evidence proves transport reachability. NHRP proves peer discovery and mapping. IPsec proves protected transport. Routing proves the destination path. CEF and forwarding evidence prove what the router will actually do with the packet.

ENGINEERING EVIDENCE · DMVPN INCIDENT LAB

INVESTIGATE — DMVPN Failure Workflow

Inject a controlled DMVPN failure, collect evidence from each dependency layer, and identify the first failed component rather than treating the final symptom as the root cause.

NO INCIDENT
Live DMVPN Incident
CONTROL PLANE + DATA PLANE
No active failure
Select a failure scenario, inject it into the topology, then investigate the dependency chain.
S1
SPOKE 1
198.51.100.11
READY
H
HUB / NHS
198.51.100.1
READY
S2
SPOKE 2
198.51.100.12
READY
FAILURE DETECTED
Engineering target: determine which dependency fails first.
Fault Injection
SELECT INCIDENT
Evidence Collection
SEQUENTIAL
1
Underlay
WAIT
2
mGRE Tunnel
WAIT
3
NHRP
WAIT
4
IPsec
WAIT
5
Routing
WAIT
6
Forwarding
WAIT
Diagnosis
ROOT CAUSE
INVESTIGATION READY

No diagnosis yet

Inject a fault and run the investigation to correlate the evidence.

R1# show dmvpn
```
☕
Support Network Insight
Help support independent engineering guides, interactive networking tools, and future learning labs.
Support Network Insight ```
eBOOK on SASE Capabilities

eBOOK – SASE Capabilities

In the following ebook, we will address the key points:

  1. Challenging Landscape
  2. The rise of SASE based on new requirements
  3. SASE definition
  4. Core SASE capabilities
  5. Final recommendations

 

 

Preliminary Information: Useful Links to Relevant Content

For pre-information, you may find the following links useful:

 

Secure Access Service Edge (SASE) is a service designed to provide secure access to cloud applications, data, and infrastructure from anywhere. It allows organizations to securely deploy applications and services from the cloud while managing users and devices from a single platform. As a result, SASE simplifies the IT landscape and reduces the cost and complexity of managing security for cloud applications and services.

SASE provides a unified security platform allowing organizations to connect users and applications securely without managing multiple security solutions. It offers secure access to cloud applications, data, and infrastructure with a single set of policies, regardless of the user’s physical location. It also enables organizations to monitor and control all user activities with the same set of security policies.

SASE also helps organizations reduce the risk of data breaches and malicious actors by providing visibility into user activity and access. In addition, it offers end-to-end encryption, secure authentication, and secure access control. It also includes threat detection, advanced analytics, and data loss prevention.

SASE allows organizations to scale their security infrastructure quickly and easily, providing them with a unified security platform that can be used to connect users and applications from anywhere securely. With SASE, organizations can quickly and securely deploy applications and services from the cloud while managing users and devices from a single platform.

 

 

 

 

POZNAN, POL - APR 15, 2021: Laptop computer displaying logo of OpenShift, a family of containerization software products developed by Red Hat

OpenShift Networking

OpenShift Networking

OpenShift, developed by Red Hat, is a leading container platform that enables organizations to streamline their application development and deployment processes. With its robust networking capabilities, OpenShift provides a secure and scalable environment for running containerized applications. This blog post will explore the critical aspects of OpenShift networking and how it can benefit your organization.

OpenShift networking is built on top of Kubernetes networking and extends its capabilities to provide a flexible and scalable networking solution for containerized applications. It offers various networking options to meet the diverse needs of different organizations.

Load balancing and service discovery are essential aspects of Openshift networking. In this section, we will explore how Openshift handles load balancing across pods using services. We will discuss the various load balancing algorithms available and highlight the importance of service discovery in ensuring seamless communication between microservices within an Openshift cluster.
<brsection 2:="" networking="" models="" in="" openshift.=""

Openshift offers different networking models to suit diverse deployment scenarios. We will explore the three main models: Overlay Networking, Host Networking, and VxLAN Networking. Each model has its advantages and considerations, and we'll highlight the use cases where they shine.

Openshift provides several advanced networking features that enhance performance, security, and flexibility. We'll dive into topics like Network Policies, Service Mesh, Ingress Controllers, and Load Balancing. Understanding and utilizing these features will empower you to optimize your Openshift networking environment.

Highlights:OpenShift Networking

**The Basics of OpenShift Networking**

OpenShift networking is built on the Kubernetes foundation, with enhancements that cater to enterprise needs. At its core, OpenShift uses the concept of Software Defined Networking (SDN) to manage container communication. This allows developers to focus on building applications rather than worrying about the complexities of network configurations. The SDN provides a virtual network layer, enabling pods to communicate without relying on the physical network infrastructure.

**Key Components and Concepts**

To navigate through OpenShift networking, it is essential to understand its key components. The primary elements include Pods, Services, Routes, and Ingress Controllers. Pods are the smallest deployable units, and they communicate with each other using the cluster’s internal IP address. Services act as a stable endpoint for accessing Pods, even as they are scaled or replaced. Routes and Ingress Controllers manage external access, allowing users to interact with the applications hosted within the cluster. These components work in harmony to provide a comprehensive networking solution.

**Security and Policies in OpenShift Networking**

Security is a paramount concern in any networking model, and OpenShift addresses this with robust security policies. Network policies in OpenShift allow administrators to control traffic flow to and from pods, enhancing the security posture of applications. By defining rules that specify which pods can communicate, administrators can create a micro-segmented network environment. This granular control helps in mitigating potential threats and reducing the attack surface.

**Advanced Networking Features**

OpenShift offers advanced networking features that cater to diverse use cases. One such feature is the Multi-Cluster Networking capability, allowing applications to span across multiple OpenShift clusters. This is particularly beneficial for organizations that operate in hybrid or multi-cloud environments. Additionally, OpenShift supports network plug-ins that extend its capabilities, such as integrating with third-party solutions for monitoring and managing network traffic.

Overview – OpenShift Networking

OpenShift networking provides a robust and scalable network infrastructure for applications running on the platform. It allows containers and pods to communicate with each other and external systems and services. The networking model in OpenShift is based on the Kubernetes networking model, providing a standardized and flexible approach.

Software-Defined Networking (SDN) is a crucial component of OpenShift Networking, enabling dynamic, programmatically efficient network configuration. SDN abstracts the traditional networking hardware, providing a flexible network fabric that can adapt to the needs of your applications. SDN operates within OpenShift, enabling improved scalability, enhanced security, and simplified network management.

Key Initial Considerations:

A. OpenShift Container Platform

OpenShift Container Platform (formerly known as OpenShift Enterprise) or OCP is Red Hat’s offering for the on-premises private platform as a service (PaaS). OpenShift is based on the Origin open-source project and is a Kubernetes distribution. The foundation of the OpenShift Container Platform is based on Kubernetes and, therefore, shares some of the same networking technology along with some enhancements.

B. Kubernetes: Orchestration Layer

Kubernetes is the leading container orchestration, and OpenShift is derived from containers, with Kubernetes as the orchestration layer. All of these elements lie upon an SDN layer that glues everything together by providing an abstraction layer. SDN creates the cluster-wide network. The glue that connects all the dots is the overlay network that operates over an underlay network. 

C. OpenShift Networking & Plugins

When it comes to container orchestration, networking plays a pivotal role in ensuring connectivity and communication between containers, pods, and services. Openshift Networking provides a robust framework that enables efficient and secure networking within the Openshift cluster. By leveraging various networking plugins and technologies, it facilitates seamless communication between applications and allows load balancing, service discovery, and more.

The Architecture of OpenShift Networking

OpenShift’s networking model is built upon the foundation of Kubernetes, but it takes things a step further with its own enhancements bringing better networking and security capabilities than the default Kubernetes model. The architecture consists of several key components, including the OpenShift SDN (Software Defined Networking), network policies, and service mesh capabilities.

The OpenShift SDN abstracts the underlying network infrastructure, allowing developers to focus on application logic rather than network configurations. This abstraction enables greater flexibility and simplifies the deployment process.

**Network Policies: Securing Your Cluster**

One of the standout features of OpenShift networking is the implementation of network policies. These policies allow administrators to define how pods communicate with each other and with the outside world. By leveraging network policies, teams can enforce security boundaries, ensuring that only authorized traffic is allowed to flow within the cluster. This is particularly crucial in multi-tenant environments where different teams might share the same OpenShift cluster but require isolated communication channels.

**Service Mesh: Enhancing Microservices Communication**

As organizations increasingly adopt microservices architectures, managing service-to-service communication becomes a challenge. OpenShift addresses this with its service mesh capabilities, primarily through the integration of Istio. A service mesh adds an additional layer of security, observability, and resilience to your microservices. It provides features like advanced traffic management, circuit breaking, and telemetry collection, all of which are essential for maintaining a healthy and efficient microservices ecosystem.

**Navigating Through OpenShift’s Route and Ingress**

Routing is another fundamental aspect of OpenShift networking. OpenShift’s routing layer allows developers to expose their applications to external users. This can be achieved through both routes and ingress resources. Routes are specific to OpenShift and provide a simple way to map external URLs to services within the cluster. On the other hand, ingress resources, which are part of Kubernetes, offer more advanced configurations and are particularly useful when you need to define complex routing rules.

Example Technology: Cloud Service Mesh

### The Core Components of a Service Mesh

At its heart, a service mesh consists of two fundamental components: the data plane and the control plane. The data plane is responsible for managing the actual data transfer between services, often through a network of proxies. These proxies intercept requests and responses, enabling smooth traffic flow and providing essential features like load balancing and retries. In contrast, the control plane oversees the configuration and management of the proxies, ensuring that policies are consistently applied across the network.

### Benefits of Implementing a Service Mesh

Implementing a cloud service mesh offers numerous advantages. Primarily, it enhances the observability of microservices by providing detailed insights into traffic patterns, latency, and service health. Additionally, it strengthens security with features like mutual TLS for encrypted communications and fine-grained access controls. Service mesh also improves resilience by offering built-in capabilities for circuit breaking and fault injection, which help to identify and mitigate potential failures proactively.

Key Considerations: OpenShift Networking

1. Network Namespace Isolation:
OpenShift networking leverages network namespaces to achieve isolation between different projects, or namespaces, on the platform. Each project has its virtual network, ensuring that containers and pods within a project can communicate securely while isolated from other projects.

2. Service Discovery and OpenShift Load Balancer:
OpenShift networking provides service discovery and load-balancing mechanisms to facilitate communication between various application components. Services act as stable endpoints, allowing containers and pods to connect to them using DNS or environmental variables. The built-in OpenShift load balancer ensures that traffic is distributed evenly across multiple instances of a service, improving scalability and reliability.

3. Ingress and Egress Network Policies:
OpenShift networking allows administrators to define ingress and egress network policies to control network traffic flow within the platform. Ingress policies specify rules for incoming traffic, allowing or denying access to specific services or pods. Egress policies, on the other hand, regulate outgoing traffic from pods, enabling administrators to restrict access to external systems or services.

4. Network Plugins and Providers:
OpenShift networking supports various network plugins and providers, allowing users to choose the networking solution that best fits their requirements. Some popular options include Open vSwitch (OVS), Flannel, Calico, and Multus. These plugins provide additional capabilities such as network isolation, advanced routing, and security features.

5. Network Monitoring and Troubleshooting:
OpenShift provides robust monitoring and troubleshooting tools to help administrators track network performance and resolve issues. The platform integrates with monitoring systems like Prometheus, allowing users to collect and analyze network metrics. Additionally, OpenShift provides logging and debugging features to aid in identifying and resolving network-related problems.

Example Technology: Prometheus Pull Approach 

**The Basics: What is the Prometheus Pull Approach?**

Prometheus utilizes a pull-based data collection model, meaning it periodically scrapes metrics from configured endpoints. This contrasts with the push model, where monitored systems send the data to a central server. The pull approach allows Prometheus to control the rate of data collection and manage the load on both the server and clients efficiently. This decoupling of data collection from data processing is a key feature that offers significant flexibility and control.

**Advantages of the Pull Model**

One of the primary advantages of the pull model is that it enables dynamic service discovery. Prometheus can automatically adjust to changes in the infrastructure, such as scaling up or down, by discovering new targets as they come online. This eliminates the need for manual configuration and reduces the risk of stale or missing data. Additionally, the pull model simplifies security configurations, as it requires fewer open ports and allows for more fine-grained access control.

Prometheus YAML file

OpenShift Networking Features

OpenShift Networking offers many features that empower developers and administrators to build and manage scalable applications. Some notable features include:

1. SDN Integration: Openshift seamlessly integrates with Software-Defined Networking (SDN) solutions, allowing for flexible network configurations and efficient traffic routing.

2. Multi-tenancy Support: With Openshift Networking, you can create isolated network zones, enabling multiple teams or projects to coexist within the same cluster while maintaining secure communication.

3. Service Load Balancing: Openshift Networking incorporates built-in load balancing capabilities, distributing incoming traffic across multiple instances of a service, thus ensuring high availability and optimal performance.

**Key Challenges: Traditional data center**

Several challenges with traditional data center networks prove they cannot support today’s applications, such as microservices and containers. Therefore, we need a new set of networking technologies built into OpenShift SDN to deal adequately with today’s landscape changes.

– Firstly, one of the main issues is that we have a tight coupling with all the networking and infrastructure components. With traditional data center networking, Layer 4 is coupled with the network topology at fixed network points and lacks the flexibility to support today’s containerized applications, which are more agile than traditional monolith applications.

– One of the main issues is that containers are short-lived and constantly spun down. Assets that support the application, such as IP addresses, firewalls, policies, and overlay networks that glue the connectivity, are continually recycled. These changes bring a lot of agility and business benefits, but there is an extensive comparison to a traditional network that is relatively static, where changes happen every few months.

OpenShift networking – Two layers

  1. In the case of OpenShift itself deployed in the virtual environment, the physical network equipment directly determines the underlying network topology. OpenShift does not control this level, which provides connectivity to OpenShift masters and nodes.
  2. OpenShift SDN plugin determines the virtual network topology. At this level, applications are connected, and external access is provided.

OpenShift uses an overlay network based on VXLAN to enable containers to communicate with each other. Layer 2 (Ethernet) frames can be transferred across Layer 3 (IP) networks using the Virtual Extensible Local Area Network (VXLAN) protocol.

Whether the communication is limited to pods within the same project or completely unrestricted depends on the SDN plugin being used. The network topology remains the same regardless of which plugin is used. OpenShift makes its internal SDN plugins available out-of-the-box and for integration with third-party SDN frameworks. 

Example Overlay Technology: VXLAN

VXLAN is a network virtualization technology that extends Layer 2 networks over a Layer 3 infrastructure. This is achieved by encapsulating Ethernet frames within UDP packets, allowing them to traverse Layer 3 networks seamlessly. The primary goal of VXLAN is to address the limitations of traditional VLANs, particularly their scalability issues.

**Advantages of VXLAN**

One of the standout features of VXLAN is its capacity to support up to 16 million unique identifiers, vastly exceeding the 4096 limit of VLANs. This scalability makes it ideal for large-scale data centers and cloud environments. Additionally, VXLAN enhances network flexibility by enabling the creation of isolated virtual networks over a shared physical infrastructure, thereby optimizing resource utilization and improving security.

VXLAN unicast mode

OpenShift Networking Plugins

### Understanding OpenShift Networking Plugins

OpenShift networking plugins are essential for managing network traffic within a cluster. They determine how pods communicate with each other and with external networks. The choice of networking plugin can significantly impact the performance and security of your applications. OpenShift supports several plugins, each with distinct features, allowing administrators to tailor their network configuration to specific needs.

### Popular Networking Plugins in OpenShift

1. **OVN-Kubernetes**: This plugin is widely used for its scalability and support for network policies. It provides a flexible, secure, and scalable network solution that is ideal for complex deployments. OVN-Kubernetes integrates seamlessly with Kubernetes, offering enhanced network isolation and simplified network management.

2. **Calico**: Known for its simplicity and efficiency, Calico provides high-performance network connectivity. It excels in environments that require fine-grained network policies and is particularly popular in microservices architectures. Calico’s support for IPv6 and integration with Kubernetes NetworkPolicy API makes it a versatile choice.

3. **Flannel**: A simpler alternative, Flannel is suitable for users who prioritize ease of setup and basic networking functionality. It uses a flat network model, which is straightforward to configure and manage, making it a good choice for smaller clusters or those in the early stages of development.

### Choosing the Right Plugin

Selecting the appropriate networking plugin for your OpenShift environment depends on several factors, including your organization’s security requirements, scalability needs, and existing infrastructure. For instance, if your primary concern is network isolation and security, OVN-Kubernetes might be the best fit. On the other hand, if you need a lightweight solution with minimal overhead, Flannel could be more suitable.

### Implementing and Managing Networking Plugins

Once you’ve chosen a networking plugin, the next step is implementation. This involves configuring the plugin within your OpenShift cluster and ensuring it aligns with your network policies. Regular monitoring and management are crucial to maintaining optimal performance and security. Utilizing OpenShift’s built-in tools and dashboards can greatly aid in this process, providing visibility into network traffic and potential bottlenecks.

Related: Before you proceed, you may find the following posts helpful for some pre-information

  1. Kubernetes Networking 101 
  2. Kubernetes Security Best Practice 
  3. Internet of Things Theory
  4. OpenStack Neutron
  5. Load Balancing
  6. ACI Cisco

OpenShift Networking

Networking Overview

1. Pod Networking: In OpenShift, containers are encapsulated within pods, the most minor deployable units. Each pod has its IP address and can communicate with other pods within the same project or across different projects. This enables seamless communication and collaboration between applications running on different pods.

2. Service Networking: OpenShift introduces the concept of services, which act as stable endpoints for accessing pods. Services provide a layer of abstraction, allowing applications to communicate with each other without worrying about the underlying infrastructure. With service networking, you can easily expose your applications to the outside world and manage traffic efficiently.

3. Ingress and Egress: OpenShift provides a robust routing infrastructure through its built-in Ingress Controller. It lets you define rules and policies for accessing your applications outside the cluster. To ensure seamless connectivity, you can easily configure routing paths, load balancing, SSL termination, and other advanced features.

4. Network Policies: OpenShift enables fine-grained control over network traffic through network policies. You can define rules to allow or deny communication between pods based on their labels and namespaces. This helps enforce security measures and isolate sensitive workloads from unauthorized access.

5. Multi-Cluster Networking: OpenShift allows you to connect multiple clusters, creating a unified networking fabric. This enables you to distribute your applications across different clusters, improving scalability and fault tolerance. OpenShift’s intuitive interface makes managing and monitoring your multi-cluster environment easy.

6. Network Policies and Security: One key aspect of Openshift Networking is its support for network policies. These policies allow administrators to define fine-grained rules and access controls, ensuring secure communication between different components within the cluster. With Openshift Networking, organizations can enforce policies restricting traffic flow, implementing encryption mechanisms, and safeguarding sensitive data from unauthorized access.

7. Service Discovery and Load Balancing: Service discovery and load balancing are crucial for maintaining high availability and optimal performance in a dynamic container environment. Openshift Networking offers robust mechanisms for service discovery, allowing containers to locate and communicate with one another seamlessly. Additionally, it provides load-balancing capabilities to distribute incoming traffic efficiently across multiple instances, ensuring optimal resource utilization and preventing bottlenecks.

8. Network Plugin Options: Openshift Networking offers a range of network plugin options, allowing organizations to select the most suitable solution for their specific requirements. Whether the native OpenShift SDN (Software-Defined Networking) plugin, the popular Flannel plugin, or the versatile Calico plugin, Openshift provides flexibility and compatibility with different network architectures and setups.

Example Network Policy Technology: GKE Kubernetes

Kubernetes network policy

POD Networking

Each pod in Kubernetes is assigned an IP address from an internal network that allows pods to communicate with each other. By doing this, all containers within the pod behave as if they were on the same host. The IP address of each pod enables pods to be treated like physical hosts or virtual machines.

It includes port allocation, networking, naming, service discovery, load balancing, application configuration, and migration. Linking pods together is unnecessary, and IP addresses shouldn’t be used to communicate directly between pods. Instead, create a service to interact with the pods.

OpenShift Container Platform DNS

To enable the frontend pods to communicate with the backend services when running multiple services, such as frontend and backend, environment variables are created for user names, service IPs, and more. To pick up the updated values for the service IP environment variable, the frontend pods must be recreated if the service is deleted and recreated.

To ensure that the IP address for the backend service is generated correctly and that it can be passed to the frontend pods as an environment variable, the backend service must be created before any frontend pods.

Due to this, the OpenShift Container Platform has a built-in DNS, enabling the service to be reached by both the service DNS and the service IP/port. Split DNS is supported by the OpenShift Container Platform by running SkyDNS on the master, which answers DNS queries for services. By default, the master listens on port 53.

POD Network & POD Communication

As a general rule, pod-to-pod communication holds for all Kubernetes clusters: An IP address is assigned to each Pod in Kubernetes. While pods can communicate directly with each other by addressing their IP addresses, it is recommended that they use Services instead. Services consist of Pods accessed through a single, fixed DNS name or IP address. The majority of Kubernetes applications use Services to communicate. Since Pods can be restarted frequently, addressing them directly by name or IP is highly brittle. Instead, use a Service to manage another pod.

Simple pod-to-pod communication

The first thing to understand is how Pods communicate within Kubernetes. Kubernetes provides IP addresses for each Pod. IP addresses are used to communicate between pods at a very primitive level. Therefore, you can directly address another Pod using its IP address whenever needed.

A Pod has the same characteristics as a virtual machine (VM), which has an IP address, exposes ports, and interacts with other VMs on the network via IP address and port.

What is the communication mechanism between the front-end pod and the back-end pod? In a web application architecture, a front-end application is expected to talk to a backend, an API, or a database. In Kubernetes, the front and back end would be separated into two Pods.

The front end could be configured to communicate directly with the back end via its IP address. However, a front end would still need to know the backend’s IP address, which can be tricky when the Pod is restarted or moved to another node. Using a Service can make our solution less brittle.

Because the app still communicates with the API pods via the Service, which has a stable IP address, if the Pods die or need to be restarted, this won’t affect the app.

pod networking
Diagram: Pod networking. Source is tutorialworks

How do containers in the same Pod communicate?

Sometimes, you may need to run multiple containers in the same Pod. The IP addresses of various containers in the same Pod are the same, so Localhost can be used to communicate between them. For example, a container in a pod can use the address localhost:8080 to communicate with another container in the Pod on port 8080.

Two containers cannot share the same port in the pod because the IP addresses are shared, and communication occurs on localhost. For instance, you wouldn’t be able to have two containers in the same Pod that expose port 8080. So, it would help if you ensured that the services use different ports.

In Kubernetes, pods can communicate with each other in a few different ways:

  1. Containers in the same Pod can connect using localhost; the other container exposes the port number.
  2. A container in a Pod can connect to another Pod using its IP address. To find the IP address of a pod, you can use oc get pods.
  3. A container can connect to another Pod through a Service. Like my service, a service has an IP address and usually a DNS name.

OpenShift and Pod Networking

When you initially deploy OpenShift, a private pod network is created. Each pod in your OpenShift cluster is assigned an IP address on the pod network, which is used to communicate with each pod across the cluster.

The pod network spanned all nodes in your cluster and was extended to your second application node when that was added to the cluster. Your pod network IP addresses can’t be used on your network by any network that OpenShift might need to communicate with. OpenShift’s internal network routing follows all the rules of any network and multiple destinations for the same IP address lead to confusion.

**Endpoint Reachability**

Also, Endpoint Reachability. Not only have endpoints changed, but have the ways we reach them. The application stack previously had very few components, maybe just a cache, web server, or database. Using a load balancing algorithm, the most common network service allows a source to reach an application endpoint or load balance to several endpoints.

A simple round-robin or a load balancer that measured load was standard. Essentially, the sole purpose of the network was to provide endpoint reachability. However, changes inside the data center are driving networks and network services toward becoming more integrated with the application.

Nowadays, the network function exists no longer solely to satisfy endpoint reachability; it is fully integrated. In the case of Red Hat’s OpenShift, the network is represented as a Software-Defined Networking (SDN) layer. SDN means different things to different vendors. So, let me clarify in terms of OpenShift.

Highlighting software-defined network (SDN)

When you examine traditional networking devices, you see the control and forwarding planes shared on a single device. The concept of SDN separates these two planes, i.e., the control and forwarding planes are decoupled. They can now reside on different devices, bringing many performance and management benefits.

The benefits of network integration and decoupling make it much easier for the applications to be divided into several microservice components driving the microservices culture of application architecture. You could say that SDN was a requirement for microservices.

software defined networking
Diagram: Software Defined Networking (SDN). Source is Opennetworking

Challenges to Docker Networking 

Port mapping and NAT

Docker containers have been around for a while, but networking had significant drawbacks when they first came out. If you examine container networking, for example, Docker containers have other challenges when they connect to a bridge on the node where the docker daemon is running.

To allow network connectivity between those containers and any endpoint external to the node, we need to do some port mapping and Network Address Translation (NAT). This adds complexity.

Port Mapping and NAT have been around for ages. Introducing these networking functions will complicate container networking when running at scale. It is perfectly fine for 3 or 4 containers, but the production network will have many more endpoints. The origins of container networking are based on a simple architecture and primarily a single-host solution.

Docker at scale: Orchestration layer

The core building blocks of containers, such as namespaces and control groups, are battle-tested. Although the docker engine manages containers by facilitating Linux Kernel resources, it’s limited to a single host operating system. Once you get past three hosts, networking is hard to manage. Everything needs to be spun up in a particular order, and consistent network connectivity and security, regardless of the mobility of the workloads, are also challenged.

Docker Default networking
Diagram: Docker Default networking

This led to an orchestration layer. Just as a container is an abstraction over the physical machine, the container orchestration framework is an abstraction over the network. This brings us to the Kubernetes networking model, which Openshift takes advantage of and enhances; for example, the OpenShift Route Construct exposes applications for external access.

The Kubernetes model: Pod networking

As we discussed, the Kubernetes networking model was developed to simplify Docker container networking, which had drawbacks. It introduced the concept of Pod and Pod networking, allowing multiple containers inside a Pod to share an IP namespace. They can communicate with each other on IPC or localhost.

Nowadays, we place a single container into a pod, which acts as a boundary layer for any cluster parameters directly affecting the container. So, we run deployment against pods rather than containers.

In OpenShift, we can assign networking and security parameters to Pods that will affect the container inside. When an app is deployed on the cluster, each Pod gets an IP assigned, and each Pod could have different applications.

For example, Pod 1 could have a web front end, and Pod could be a database, so the Pods need to communicate. For this, we need a network and IP address. By default, Kubernetes allocates an internal IP address for each Pod for applications running within the Pod. Pods and their containers can network, but clients outside the cluster cannot access internal cluster resources by default. With Pod networking, every Pod must be able to communicate with each other Pod in the cluster without Network Address Translation (NAT).

A typical service type: ClusterIP

The most common service IP address type is “ClusterIP .” ClusterIP is a persistent virtual IP address used for load-balancing traffic internal to the cluster. Services with these service types cannot be directly accessed outside the cluster; there are other service types for that requirement.

The service type of Cluster-IP is considered for East-West traffic since it originates from Pods running in the cluster to the service IP backed by Pods that also run in the cluster.

Then, to enable external access to the cluster, we need to expose the services that the Pod or Pods represent, and this is done with an Openshift Route that provides a URL. So, we have a service in front of the pod or groups of pods. The default is for internal access only. Then, we have a URL-based route that gives the internal service external access.

Openshift load balancer
Diagram: Openshift networking and clusterIP. Source is Redhat.

Using an OpenShift Load Balancer

Get Traffic into the Cluster

If you do not need a specific external IP address, OpenShift Container Platform clusters can be accessed externally through an OpenShift load balancer service. The OpenShift load balancer allocates unique IP addresses from configured pools. Load balancers have a single edge router IP (which can be a virtual IP (VIP), but it is still a single machine for initial load balancing). How many OpenShift load balancers are there in OpenShift?

Two load balancers

The solution supports some load balancer configuration options: Use the playbooks to configure two load balancers for highly available production deployments, or use the playbooks to configure a single load balancer, which is helpful for proof-of-concept deployments. Deploy the solution using your OpenShift load balancer.

This process involves the following:

  1. The administrator performs the prerequisites;
  2. The developer creates a project and service if the service to be exposed does not exist;
  3. The developer exposes the service to create a route.
  4. The developer creates the Load Balancer Service.
  5. The network administrator configures networking to the service.

OpenShift load balancer: Different Openshift SDN networking modes

OpenShift security best practices  

So, depending on your Openshift SDN configuration, you can tailor the network topology differently. You can have free-for-all Pod connectivity, similar to a flat network or something stricter, with different security boundaries and restrictions. A free-for-all Pod connectivity between all projects might be good for a lab environment.

Still, you may need to tailor the network with segmentation for production networks with multiple projects, which can be done with one of the OpenShift SDN plugins. We will get to this in a moment.

Openshift networking does this with an SDN layer and enhances Kubernetes networking to have a virtual network across all the nodes created with the Open switch standard. For the Openshift SDN, this Pod network is established and maintained by the OpenShift SDN, configuring an overlay network using Open vSwitch (OVS).

The OpenShift SDN plugin

We mentioned that you could tailor the virtual network topology to suit your networking requirements. The OpenShift SDN plugin and the SDN model you select can determine this. With the default OpenShift SDN, several modes are available.

This level of SDN mode you choose is concerned with managing connectivity between applications and providing external access to them.

Some modes are more fine-grained than others. How are all these plugins enabled? The Openshift Container Platform (OCP) networking relies on the Kubernetes CNI model while supporting several plugins by default and several commercial SDN implementations, including Cisco ACI.

The native plugins rely on the virtual switch Open vSwitch and offer alternatives to providing segmentation using VXLAN, specifically the VNID or the Kubernetes Network Policy objects:

We have, for example:

        • ovs-subnet  

        • ovs-multitenant  

        • ovs-network policy

Choosing the right plugin depends on your security and control goals. As SDNs take over networking, third-party vendors develop programmable network solutions. OpenShift is tightly integrated with products from such providers by Red Hat. According to Red Hat, the following solutions are production-ready:

  1. Nokia Nuage
  2. Cisco Contiv
  3. Juniper Contrail
  4. Tigera Calico
  5. VMWare NSX-T
  6. ovs-subnet plugin

After OpenShift is installed, this plugin is enabled by default. As a result, pods can be connected across the entire cluster without limitations so traffic can flow freely between them. This may be undesirable if security is a top priority in large multitenant environments. 

  • ovs-multitenant plugin

Security is usually unimportant in PoCs and sandboxes but becomes paramount when large enterprises have diverse teams and project portfolios, especially when third parties develop specific applications. A multitenant plugin like ovs-multitenant is an excellent choice if simply separating projects is all you need.

This plugin sets up flow rules on the br0 bridge to ensure that only traffic between pods with the same VNID is permitted, unlike the ovs-subnet plugin, which passes all traffic across all pods. It also assigns the same VNID to all pods for each project, keeping them unique across projects.

  • ovs-networkpolicy plugin

While the ovs-multitenant plugin provides a simple and largely adequate means for managing access between projects, it does not allow granular control over access. In this case, the ovs-networkpolicy plugin can be used to create custom NetworkPolicy objects that, for example, apply restrictions to traffic egressing or entering the network.

  • Egress routers

In OpenShift, routers direct ingress traffic from external clients to services, which then forward it to pods. As well as forwarding egress traffic from pods to external networks, OpenShift offers a reverse type of router. Egress routers, on the other hand, are implemented using Squid instead of HAProxy. Routers with egress capabilities can be helpful in the following situations:

They are masking different external resources used by several applications with a single global resource. For example, applications may be developed so that they are built, pulling dependencies from other mirrors, and collaboration between their development teams is rather loose. So, instead of getting them to use the same mirror, an operations team can set up an egress router to intercept all traffic directed to those mirrors and redirect it to the same site.

To redirect all suspicious requests for specific sites to the audit system for further analysis.

OpenShift supports the following types of egress routers:

  • redirect for redirecting traffic to a specific destination IP
  • http-proxy for proxying HTTP, HTTPS, and DNS traffic

Summary:OpenShift Networking

In the ever-evolving world of cloud computing, Openshift has emerged as a robust application development and deployment platform. One crucial aspect that makes it stand out is its networking capabilities. In this blog post, we delved into the intricacies of Openshift networking, exploring its key components, features, and benefits.

Understanding Openshift Networking Fundamentals

Openshift networking operates on a robust and flexible architecture that enables efficient communication between various components within a cluster. It utilizes a combination of software-defined networking (SDN) and network overlays to create a scalable and resilient network infrastructure.

Exploring Networking Models in Openshift

Openshift offers different networking models to suit various deployment scenarios. The most common models include Single-Stacked Networking, Dual-Stacked Networking, and Multus CNI. Each model has advantages and considerations, allowing administrators to choose the most suitable option for their specific requirements.

Deep Dive into Openshift SDN

At the core of Openshift networking lies the Software-Defined Networking (SDN) solution. It provides the necessary tools and mechanisms to manage network traffic, implement security policies, and enable efficient communication between Pods and Services. We will explore the inner workings of Openshift SDN, including its components like the SDN controller, virtual Ethernet bridges, and IP routing.

Network Policies in Openshift

To ensure secure and controlled communication between Pods, Openshift implements Network Policies. These policies define rules and regulations for network traffic, allowing administrators to enforce fine-grained access controls and segmentation. We will discuss the concept of Network Policies, their syntax, and practical examples to showcase their effectiveness.

Conclusion: Openshift’s networking capabilities play a crucial role in enabling seamless communication and connectivity within a cluster. By understanding the fundamentals, exploring different networking models, and harnessing the power of SDN and Network Policies, administrators can leverage Openshift’s networking features to build robust and scalable applications.

In conclusion, Openshift networking opens up a world of possibilities for developers and administrators, empowering them to create resilient and interconnected environments. By diving deep into its intricacies, one can unlock the full potential of Openshift networking and maximize the efficiency of their applications.

auto scaling observability

Auto Scaling Observability

Autoscaling Observability

Modern applications rarely operate at a constant workload. Traffic can change rapidly, services can be distributed across multiple hosts, and workloads may scale up or down automatically in response to demand. In this environment, autoscaling observability provides the visibility required to understand not only whether a system is scaling, but whether scaling decisions are producing the desired performance, reliability, and resource-efficiency outcomes.

Autoscaling is fundamentally a control loop. The platform observes one or more signals, compares them against a defined target or threshold, and adjusts capacity accordingly. These signals might include CPU utilisation, memory consumption, request rate, queue depth, latency, or application-specific metrics. Observability provides the surrounding evidence needed to determine whether those signals accurately represent workload demand and whether the resulting scaling behaviour is effective.

The foundation is the collection of metrics, logs, and traces. Metrics provide continuously measurable signals such as request rate, saturation, resource utilisation, replica count, scaling events, and response latency. Logs provide contextual information about application and infrastructure behaviour, while distributed traces allow individual requests to be followed across multiple services. Together, these telemetry sources provide a much more complete view of what is happening during scaling events.

Metrics and alerting are particularly important because autoscaling can introduce behaviours that are not immediately visible from infrastructure utilisation alone. An application may scale because CPU usage increases while user-facing latency continues to deteriorate, or it may repeatedly scale up and down because the selected metric is too sensitive. Monitoring replica counts, scaling events, request rates, saturation, latency, and error rates helps operators identify ineffective scaling policies and unstable control loops.

Logs and event data provide the operational context behind those metrics. Scaling events, failed deployments, application errors, resource constraints, scheduling problems, and configuration changes can explain why a workload behaved differently during a particular period. Correlating these events with metric data makes it easier to distinguish genuine demand changes from infrastructure or application problems.

For distributed applications, distributed tracing adds another dimension of visibility. A service may appear healthy at the infrastructure level while downstream dependencies introduce significant latency. Tracing individual requests across services can reveal where latency or failures occur during periods of increased demand and whether additional replicas are actually addressing the underlying bottleneck.

Effective autoscaling observability also requires carefully selected objectives. Rather than collecting every available metric, teams should identify the signals that represent user experience, workload demand, resource saturation, and system capacity. These signals can then be connected to dashboards, alerts, scaling policies, and operational runbooks.

Tool selection is equally important. A modern observability architecture may combine a metrics platform such as Prometheus, log aggregation, distributed tracing, dashboarding, cloud-native monitoring services, and application instrumentation. The objective is not simply to accumulate telemetry, but to create a coherent view in which scaling decisions can be analysed against application behaviour and business impact.

Finally, automation should extend beyond the scaling mechanism itself. Automated alerting, anomaly detection, telemetry collection, and event correlation can reduce the time required to identify scaling problems. However, automation should remain observable and explainable: engineers need to understand what caused a scaling decision, what changed afterwards, and whether the resulting behaviour improved the system.

Autoscaling observability therefore sits at the intersection of monitoring, application performance, infrastructure capacity, and automated control. By correlating scaling signals with metrics, logs, traces, and user-facing outcomes, organisations can move beyond simply adding or removing capacity and instead build systems that scale predictably, efficiently, and with measurable reliability.

Auto-Scaling Decision Lab

Change the telemetry signals and see how a simplified autoscaling engine decides whether to scale out, hold capacity, or scale in.

1. Adjust Observability Signals

Request Rate Normal
LowNormalHighSurge
CPU Utilisation Normal
30%55%75%90%
p95 Latency Normal
80ms150ms300ms700ms
Queue Depth Normal
LowNormalHighCritical
Error Rate Normal
0.1%0.5%2%5%+

2. Autoscaling Decision Path

TelemetrySignals
→
PolicyTargets
→
AutoscalerDecision
CapacityPods / Nodes
←
ActionScale
←
FeedbackNew telemetry
HOLD CAPACITY
Signals are within the expected operating range. No immediate capacity change is required.
Primary driverBalanced workload
Recommended actionMaintain capacity
Engineering takeaway: Autoscaling decisions are driven by configured metrics and policies. CPU may be the target for an HPA, while application metrics such as queue depth or request rate can provide a better signal for some workloads. Real autoscalers also use cooldowns, stabilization windows and capacity limits; this lab simplifies those mechanisms so the relationship between telemetry and scaling decisions is visible.

Highlights: Auto-scaling Observability

Understanding Auto-Scaling Observability

1: ) Auto-scaling observability refers to a system’s ability to adjust its resources based on real-time monitoring and analysis automatically. It combines two essential components: auto-scaling, which dynamically allocates resources, and observability, which provides insights into the system’s behavior and performance.

2: ) By leveraging advanced monitoring tools and intelligent algorithms, auto-scaling observability enables organizations to optimize resource allocation and respond swiftly to changing demands.

3: ) Auto scaling observability involves the continuous monitoring and analysis of system metrics to automate the scaling of resources. This process allows organizations to automatically increase or decrease their computing power based on current demand, ensuring that applications run smoothly without over-provisioning or incurring unnecessary costs.

Auto‑Scaling Observability

Auto‑scaling observability is the ability to see, measure, and understand how your infrastructure scales up and down in real time — across containers, microservices, Kubernetes clusters, serverless functions, cloud workloads, and distributed systems. It ensures that scaling decisions are correct, efficient, cost‑optimized, and aligned with application performance and user experience.

This section is written cleanly for direct copy‑and‑paste into your architecture notes or blog.

Core Observability Signals for Auto‑Scaling

  • Metrics — CPU, memory, queue depth, request rate, latency, p95/p99 spikes.

  • Logs — scaling events, pod lifecycle logs, autoscaler decisions.

  • Traces — request flow timing to detect bottlenecks before scaling triggers.

  • Kubernetes Events — HPA/VPA decisions, pod scheduling, node pressure.

  • Cloud Telemetry — AWS ASG, Azure VMSS, GCP MIG scaling signals.

  • Application Performance — throughput, error rate, saturation, concurrency.

These signals allow you to understand why scaling happened, whether it was correct, and how it impacted performance.

What Auto‑Scaling Observability Provides

1. Visibility into Scaling Decisions

You see exactly why the system scaled:

  • CPU threshold exceeded

  • latency spike

  • queue backlog

  • memory pressure

  • traffic surge

  • degraded user experience

2. Real‑Time Performance Correlation

Scaling is correlated with:

  • request rate

  • p95/p99 latency

  • error bursts

  • cold‑start delays

  • pod/container startup times

3. Cost and Efficiency Insights

Observability reveals:

  • over‑scaling

  • under‑scaling

  • wasted compute

  • inefficient autoscaler thresholds

  • unnecessary node expansions

4. Predictive Scaling Signals

Using ML‑driven analytics:

  • forecast traffic

  • detect anomalies

  • pre‑scale before peak load

  • avoid cascading failures

Auto‑Scaling Observability in Modern Architectures

Kubernetes (HPA/VPA)

Observability tracks:

  • pod lifecycle

  • node pressure

  • HPA decisions

  • VPA recommendations

  • cluster autoscaler behavior

  • container cold‑start impact

Serverless (Lambda, Functions, Cloud Run)

Observability shows:

  • concurrency spikes

  • cold‑start latency

  • throttling events

  • function execution time

  • scaling bursts

Microservices

Observability correlates:

  • service‑to‑service latency

  • dependency bottlenecks

  • trace‑driven scaling triggers

  • queue depth and backpressure

Cloud Auto‑Scaling Groups

Tracks:

  • VM provisioning time

  • scaling cooldown windows

  • load balancer health

  • network saturation

Tools for Auto‑Scaling Observability

  • Prometheus — metrics for HPA/VPA.

  • Grafana — dashboards for scaling behavior.

  • OpenTelemetry — traces + metrics for scaling triggers.

  • ThousandEyes — user experience impact during scaling events.

  • Splunk — logs + analytics + correlation.

  • CloudWatch / Azure Monitor / GCP Cloud Monitoring — native cloud scaling telemetry.

Together, these tools create a complete scaling observability fabric.

Outcome

Auto‑scaling observability delivers:

  • predictable scaling

  • reduced latency during load spikes

  • optimized resource usage

  • lower cloud cost

  • faster incident response

  • improved user experience

  • reliable microservices and Kubernetes operations

It transforms auto‑scaling from a reactive mechanism into an intelligent, data‑driven, self‑optimizing system.

Auto-Scaling – Key Points:

– Enhanced Scalability and Performance: Auto-scaling observability allows systems to scale resources up or down based on actual usage patterns. This ensures that the system can handle peak loads efficiently without overprovisioning resources during periods of low demand. Organizations can avoid costly downtime by dynamically adjusting resources and ensuring optimal performance during sudden traffic spikes.

– Cost Optimization: With auto-scaling observability, businesses can significantly reduce infrastructure costs. Organizations can avoid unnecessary idle resource expenditures by accurately provisioning resources based on real-time data. This cost optimization approach ensures that companies only pay for the resources required, resulting in considerable savings.

– Improved Fault Tolerance: Auto-scaling observability is crucial in enhancing system resilience. Organizations can promptly identify and address potential issues by continuously monitoring the system’s health. In case of anomalies or failures, the system can automatically scale resources or trigger alerts for immediate remediation. This proactive approach minimizes the impact of failures and enhances the system’s overall fault tolerance.

Auto-scaling Example: Scaling with Docker Swarm

Auto Scaling – Key Components:

To fully harness the power of auto scaling observability, it’s important to understand its key components. These include metrics collection, alerting, and automated responses. Let’s delve deeper into each of these elements:

1. **Metrics Collection:** Gathering data from various sources is the foundation of observability. This involves collecting CPU usage, memory utilization, network traffic, and other vital metrics. With comprehensive data, organizations can better understand their infrastructure’s behavior and make informed scaling decisions.

2. **Alerting:** Once data is collected, it’s essential to set up alerts for any anomalies or thresholds that are breached. Alerting enables teams to respond swiftly to potential issues, minimizing downtime and maintaining application performance.

3. **Automated Responses:** The ultimate goal of observability is to automate responses to fluctuating demands. By employing pre-defined rules and machine learning algorithms, businesses can ensure that resources are scaled up or down automatically, optimizing both performance and cost.

**The Role of the Metric**

“What Is a Metric: Good for Known” Regarding auto-scaling observability and metrics, one must understand the metric’s downfall. A metric is a single number, with tags optionally appended for grouping and searching those numbers. They are disposable and cheap and have a predictable storage footprint.

A metric is a numerical representation of a system state over a recorded time interval. It can tell you if a particular resource is over or underutilized at a specific moment. For example, CPU utilization might be at 75% right now.

Implementing Auto-Scaling Observability

Choosing the Right Monitoring Tools: To effectively implement auto-scaling observability, organizations must select appropriate monitoring tools that provide real-time insights into system performance, resource utilization, and user behavior. These tools should offer robust analytics capabilities and seamless integration with auto-scaling platforms.

Defining Metrics and Thresholds: Accurate metrics and thresholds are critical for successful auto-scaling observability. Organizations must identify key performance indicators (KPIs) that align with their business objectives and set appropriate thresholds for scaling actions. For example, CPU utilization, response time, and error rates are standard metrics for auto-scaling decisions.

Automating Scaling Actions: Organizations should automate scaling actions based on predefined rules to fully leverage the benefits of auto-scaling observability. By integrating monitoring tools, auto-scaling platforms, and orchestration frameworks, businesses can ensure that resource allocation adjustments are performed seamlessly and without human intervention.

Service Mesh & Auto-Scaling

Service mesh acts as a dedicated infrastructure layer for managing service-to-service communications. It provides a suite of capabilities, including traffic management, security, and, most importantly, observability. By integrating a service mesh, such as Istio or Linkerd, into your auto scaling environment, you gain granular visibility into your microservices architecture. This includes detailed metrics, tracing, and logging, enabling you to monitor traffic patterns, latency, and error rates with precision.

### Implementing Service Mesh for Optimal Observability

Deploying a service mesh involves several considerations to maximize its observability benefits. Start by identifying the microservices that will benefit most from enhanced observability. Next, configure the service mesh to collect and process telemetry data effectively. Ensure your observability stack—comprising metrics, logs, and traces—is equipped to handle the data influx. Finally, leverage the insights gained to optimize your auto scaling strategy, ensuring minimal downtime and optimal performance.

### What is a Cloud Service Mesh?

A cloud service mesh is a dedicated infrastructure layer that manages service-to-service communication within a distributed application. It decouples the networking logic from the application code, enabling developers to focus on core functionality without worrying about the complexities of inter-service communication. Service meshes provide features like load balancing, service discovery, and security policies, making them indispensable for modern cloud-native applications.

### Key Benefits of Service Mesh

#### Simplified Networking

One of the primary benefits of a service mesh is the simplification of networking within a microservices architecture. By abstracting the communication logic, service meshes make it easier to manage and scale applications. Developers can implement features like retries, timeouts, and circuit breakers without modifying their application code.

#### Enhanced Security

Service meshes provide robust security features, including mutual TLS (mTLS) for service-to-service encryption and authentication. This ensures that communication between services is secure by default, reducing the risk of data breaches and unauthorized access.

#### Traffic Management

With a service mesh, you can intelligently route traffic between services based on various criteria such as load, service version, or geographic location. This level of control enables canary deployments, blue-green deployments, and A/B testing, making it easier to roll out new features with minimal risk.

### The Role of Observability

Observability is the ability to measure the internal states of a system based on the outputs it produces. In the context of a service mesh, observability involves collecting and analyzing metrics, logs, and traces to gain insight into the performance and behavior of the services.

### Why Observability Matters

Without proper observability, managing a service mesh can become a daunting task. Observability allows you to monitor the health of your services, detect anomalies, and troubleshoot issues in real-time. It provides the visibility needed to ensure that your service mesh is functioning as intended and that any problems are quickly identified and resolved.

### Tools and Techniques

Several tools can enhance observability in a service mesh, such as Prometheus for metrics, Jaeger for tracing, and Fluentd for logging. Combining these tools provides a comprehensive view of your service mesh’s performance and health, enabling proactive maintenance and quicker issue resolution.

Gaining Visibility with Google Ops Agent

**The Role of Google Ops Agent in Observability**

Google Ops Agent is a unified agent that simplifies the process of collecting telemetry data from your cloud environment. Its ability to seamlessly integrate with Google Cloud’s operations suite makes it an essential tool for businesses looking to enhance their auto-scaling observability. By providing detailed insights into CPU usage, memory utilization, and network traffic, Google Ops Agent ensures that your infrastructure is both monitored and optimized in real-time.

**Benefits of Enhanced Observability**

With Google Ops Agent in place, businesses gain a clearer picture of how their auto-scaling systems are functioning. This enhanced observability translates to several benefits: improved system reliability, faster troubleshooting, and better resource allocation. By actively monitoring metrics and logs, organizations can preemptively address issues before they escalate into significant problems. Furthermore, insights derived from this data can inform future scaling strategies and infrastructure investments.

**Example: Understanding Ops Agent**

Ops Agent is a lightweight and efficient monitoring agent developed by Google Cloud. It enables you to collect crucial metrics and logs from your Compute Engine instances, providing valuable insights into their performance and health. By leveraging Ops Agent, you can proactively detect issues, troubleshoot problems, and optimize the utilization of your instances.

To begin monitoring your Compute Engine instances with Ops Agent, install it on your virtual machines. The installation process is straightforward and can be done using package managers like apt or yum. Once installed, Ops Agent seamlessly integrates with Google Cloud Monitoring, allowing you to access and analyze the gathered data.

After installing Ops Agent, it is essential to configure the monitoring metrics that you want to collect. Ops Agent supports many metrics, including CPU usage, memory utilization, disk I/O, and network traffic. By tailoring the metrics collection to your specific needs, you can efficiently monitor the performance of your Compute Engine instances and identify any anomalies or bottlenecks.

What is GKE-Native Monitoring?

GKE-Native Monitoring is a powerful monitoring solution provided by Google Cloud Platform (GCP) designed explicitly for GKE clusters. It leverages the capabilities of Prometheus and Stackdriver, offering a unified monitoring experience within the GCP ecosystem. With GKE-Native Monitoring, users can effortlessly collect, visualize, and analyze metrics and logs related to their GKE clusters, enabling them to make data-driven decisions and proactively address any issues that may arise.

GKE-Native Monitoring offers a range of features that enhance observability and simplify monitoring workflows. Some notable features include:

1. Automatic Metric Collection: GKE-Native Monitoring automatically collects a rich set of metrics from every GKE cluster, including CPU and memory utilization, network traffic, and application-specific metrics. This eliminates the need for manual configuration and ensures comprehensive monitoring that is out of the box.

2. Custom Metrics and Alerts: Users can define custom metrics and alerts tailored to their applications and business requirements. This empowers them to monitor critical aspects of their clusters and receive notifications when predefined thresholds are crossed, enabling timely actions and proactive troubleshooting.

3. Integration with Stackdriver Logging: GKE-Native Monitoring integrates with Stackdriver Logging, allowing users to correlate log data with metrics. By combining logs and metrics, users can gain a holistic view of their application’s behavior and quickly identify the root causes of any issues.

Kubernetes Autoscaling

Kubernetes auto scaling involves dynamically adjusting the number of running pods in a cluster based on current demand. This ensures that applications remain responsive while optimizing resource utilization. With the right configuration, auto scaling can help maintain performance during traffic spikes and reduce costs during low-traffic periods.

**Benefits of Implementing Auto Scaling**

The implementation of Kubernetes auto scaling brings a multitude of benefits to organizations. Firstly, it enhances resource efficiency by ensuring that applications use only the necessary resources, reducing wastage and lowering operational costs. Secondly, it improves application performance and reliability by adapting to traffic fluctuations, ensuring a consistent user experience even during peak demand. Moreover, auto scaling supports rapid scaling for new deployments, enabling businesses to respond swiftly to market changes without manual intervention.

**Challenges and Best Practices**

While Kubernetes auto scaling offers significant advantages, it also presents challenges that organizations need to navigate. One common issue is configuring the right scaling metrics and thresholds to avoid over-provisioning or under-provisioning resources. It’s crucial to thoroughly test and monitor these settings in a staging environment before deploying them to production. Additionally, consider using custom metrics that align closely with your application’s performance indicators for more accurate scaling decisions. Regularly reviewing and updating your scaling policies ensures they remain effective as application workloads evolve.

### Types of Auto Scaling in Kubernetes

Kubernetes offers several types of auto scaling, each designed to address different aspects of resource management. The most commonly used are Horizontal Pod Autoscaler (HPA), Vertical Pod Autoscaler (VPA), and Cluster Autoscaler.

1. **Horizontal Pod Autoscaler (HPA):** HPA adjusts the number of pod replicas in a deployment or replication controller based on observed CPU utilization or other select metrics. It’s ideal for applications with varying workloads, ensuring that resources are available when needed and conserved when demand is low.

2. **Vertical Pod Autoscaler (VPA):** VPA automatically adjusts the CPU and memory requests and limits for containers within pods. This ensures that each pod has the right amount of resources, preventing over-provisioning and under-provisioning.

3. **Cluster Autoscaler:** This tool automatically adjusts the size of the Kubernetes cluster so that all pods have a place to run. It adds nodes when pods are unschedulable due to resource shortages and removes nodes when they’re underutilized.

### Configuring Auto Scaling in Kubernetes

To leverage Kubernetes auto scaling effectively, you’ll need to configure it to meet your application’s specific needs. The process typically involves setting up metrics and thresholds that trigger scaling actions.

For HPA, you’ll define the target CPU utilization or other custom metrics that the autoscaler should monitor. VPA requires setting up recommendations for resource requests and limits. Finally, the Cluster Autoscaler needs to be linked with your cloud provider to manage node scaling efficiently.

It’s crucial to regularly monitor and adjust these configurations to ensure optimal performance, as application demands can evolve over time.

### Best Practices for Kubernetes Auto Scaling

Implementing auto scaling in Kubernetes is not a set-it-and-forget-it task. Here are some best practices to consider:

– **Understand Your Workloads:** Analyze your application’s workload patterns to choose the right type of auto scaling and set appropriate thresholds.

– **Use Custom Metrics:** While CPU and memory are common metrics, consider using application-specific metrics to drive more accurate scaling decisions.

– **Test and Monitor:** Regularly test your auto scaling configurations in a non-production environment. Continuous monitoring and logging are essential to catch and resolve issues early.

You can dynamically scale up or down any architecture component through autoscaling.

An example of a good use of autoscaling is as follows:

You may need additional web servers to handle the surge in traffic at the end of the day when your website’s load increases. Where does the rest of the day fit in? Your servers can’t sit idle during most business hours. Especially if you’re using a cloud provider, you want to optimize your environment’s potential costs. An autoscale allows you to increase the number of components during a spike and scale down during a regular period.

Example: Prometheus Pull Approach

There can be many tools to gather metrics, such as Prometheus, along with several techniques used to collect these metrics, such as the PUSH and PULL approaches. There are pros and cons to each method. However, Prometheus metric types and its PULL approach are prevalent in the market. However, if you want full observability and controllability, remember it is solely in metrics-based monitoring solutions.  For additional information on Monitoring and Observability and their difference, visit this post on observability vs monitoring.

These autoscalers rely on Kubernetes’ metric server for scaling up or down k8s objects.

Adopting Auto-scaling

Autoscaling is a mechanism that automatically adjusts the number of computing resources allocated to an application based on its demand. By dynamically scaling resources up or down, autoscaling enables organizations to handle fluctuating workloads efficiently. However, robust observability is crucial to harness the power of autoscaling truly.

The Role of Observability in Autoscaling

Observability is the ability to gain insights into a system’s internal state based on its external outputs. It plays a pivotal role in understanding the system’s behavior, identifying bottlenecks, and making informed scaling decisions regarding autoscaling. It provides visibility into key metrics like CPU utilization, memory usage, and network traffic. With observability, you can make data-driven decisions and ensure optimal resource allocation.

 Monitoring and Metrics

To achieve effective autoscaling observability, comprehensive monitoring is essential. Monitoring tools collect various metrics, such as response times, error rates, and resource utilization, to provide a holistic view of your infrastructure. These metrics can be analyzed to identify patterns, detect anomalies, and trigger autoscaling actions when necessary. You can proactively address performance issues and optimize resource utilization by monitoring and analyzing metrics.

Logging and Tracing

In addition to monitoring, logging, and tracing are critical components of autoscaling observability. Logging captures detailed information about system events, errors, and activities, enabling you to troubleshoot issues and gain insights into system behavior. Tracing helps you understand the flow of requests across different services. Logging and tracing provide a granular view of your application’s performance, aiding in autoscaling decisions and ensuring smooth operation.

Automation and Alerting

Automation and alerting mechanisms are vital to mastering autoscaling observability. You can configure thresholds and triggers that initiate autoscaling actions based on predefined conditions by setting up automated processes. This allows for proactive scaling, ensuring your system is constantly optimized for performance. Additionally, timely alerts can notify you of critical events or anomalies, enabling you to take immediate action and maintain the desired scalability.

Autoscaling observability is the key to unlocking its true potential. By understanding your system’s behavior through comprehensive monitoring, logging, and tracing, you can make informed decisions and ensure optimal resource allocation. With automation and alerting mechanisms, you can proactively respond to changing demands and maintain high efficiency. Embrace autoscaling observability and take your infrastructure management to new heights.

Managed Instance Groups

### Auto Scaling: Adapting to Your Needs

One of the standout features of managed instance groups is auto scaling. With auto scaling, your infrastructure can dynamically adjust to the current demand. This ensures that your applications have the necessary resources without overspending. By setting up policies based on CPU usage, requests per second, or custom metrics, MIGs can efficiently allocate resources, keeping your applications responsive and your costs under control.

—

### Observability: Keeping a Close Watch

Observability is key in maintaining the health of your cloud infrastructure. Google Cloud’s managed instance groups provide comprehensive monitoring tools that give you insights into the performance and stability of your instances. By leveraging metrics, logs, and traces, you can detect anomalies, optimize performance, and ensure your applications run smoothly. This proactive approach to monitoring allows you to address potential issues before they impact your services.

—

### Integration with Google Cloud

Managed instance groups seamlessly integrate with various Google Cloud services, enhancing their utility and flexibility. From load balancing to deploying containerized applications with Google Kubernetes Engine, MIGs work in tandem with Google’s ecosystem to provide a cohesive and powerful cloud solution. This integration not only simplifies management but also boosts the scalability and reliability of your applications.

Managed Instance Group

 

Scaling Incident Investigation

The autoscaler reacted correctly — but the application is still slow. Investigate the telemetry to find the bottleneck.

INCIDENT: Checkout traffic has increased by 2.4×. HPA increased the API deployment from 4 to 10 pods, but customer-facing latency continues to rise. What is actually limiting performance?

Metrics — What is changing?

Investigation Clues

Scaling response HPA increased API replicas from 4 → 10.
API CPU CPU falls after additional pods are added.
Customer latency p95 remains above 900 ms.
Downstream dependency Database calls are consuming most request time.
Engineering conclusion: Scaling the API tier increased application capacity, but it did not remove the downstream bottleneck. Observability connects the scaling event with the user-facing symptom and shows where additional capacity stops helping.

Autoscaling Observability

Understanding Autoscaling

-Before we discuss observability, let’s briefly explore the concept of autoscaling. Autoscaling refers to the ability of an application or infrastructure to automatically adjust its resources based on demand. It enables organizations to handle fluctuating workloads and optimize resource allocation efficiently.

-Observability, in the context of autoscaling, refers to gaining insights into an autoscaling system’s performance, health, and efficiency. It involves collecting, analyzing, and visualizing relevant data to understand the application and infrastructure’s behavior and patterns.

-Through observability, organizations can make informed decisions to optimize autoscaling algorithms, resource allocation, and overall system performance. To achieve effective autoscaling observability, several critical components come into play. These include:

A. Metrics and Monitoring: Gathering and monitoring key metrics such as CPU utilization, response times, request rates, and error rates is fundamental for understanding the application and infrastructure’s performance.

B. Logging and Tracing: Logging captures detailed information about events and transactions within the system, while tracing provides insights into the flow of requests across various components. Both logging and tracing contribute to a comprehensive understanding of system behavior.

C. Alerting and Thresholds: Setting up appropriate alerts and thresholds based on predefined criteria ensures timely notifications when specific conditions are met.  

Tools and Technologies for Autoscaling Observability

A wide range of tools and technologies are available to facilitate autoscaling observability. Prominent examples include Prometheus, Grafana, Elasticsearch, Kibana, and CloudWatch. These tools provide robust monitoring, visualization, and analysis capabilities, enabling organizations to gain deep insights into their autoscaling systems.

The first component of observability is the channels that convey observations to the observer. There are three channels: logs, traces, and metrics. These channels are common to all areas of observability, including data observability.

1.Logs: Logs are the most typical channel and take several forms (e.g., a line of free text or JSON). They are intended to encapsulate information about an event.

2.Traces: Traces allow you to do what logs don’t—reconnect the dots of a process. Because traces represent the link between all events of the same process, they allow the whole context to be derived from logs efficiently. Each pair of events, an operation, is a span that can be distributed across multiple servers.

3.Metrics: Finally, we have metrics. Every system state has some component that can be represented with numbers, and these numbers change as the state changes.

Understanding VPC Flow Logs

VPC Flow Logs capture information about the IP traffic going in and out of Virtual Private Clouds (VPCs) within Google Cloud. Enabling VPC Flow Logs allows you to gain visibility into network traffic at the subnet level, thereby facilitating network troubleshooting, security analysis, and performance monitoring.

Once the VPC Flow Logs are enabled and data starts flowing in, it’s time to tap into the potential of Google Cloud Logging. Using the appropriate filters and queries, you can sift through the vast amount of log data and extract meaningful insights. Whether it’s identifying suspicious traffic patterns, monitoring network performance metrics, or investigating security incidents, Google Cloud Logging provides a robust set of tools to facilitate these analyses.

Auto Scaling Observability

**Metrics: Resource Utilization Only**

– Metrics help us understand resource utilization. In a Kubernetes environment, these metrics are used for auto-healing and auto-scheduling. Monitoring performs several functions when it comes to metrics. First, it can collect, aggregate, and analyze metrics to identify known patterns that indicate troubling trends.

– The critical point here is that it shifts through known patterns. Then, based on a known event, metrics trigger alerts that notify when further investigation is needed. Finally, we have dashboards that display the metrics data trends adapted for visual consumption.

– These monitoring systems work well for identifying previously encountered known failures but don’t help as much for the unknown. Unknown failures are the norm today with disgruntled systems and complex system interactions.

– Metrics are suitable for dashboards, but there won’t be a predefined dashboard for unknowns, as it can’t track something it does not know about. Using metrics and dashboards like this is a reactive approach, yet it’s widely accepted as the norm. Monitoring is a reactive approach best suited for detecting known problems and previously identified patterns. 

**Metrics and intermittent problems**

– The metrics can help you determine whether a microservice is healthy or unhealthy within a microservices environment. Still, a metric will have difficulty telling you if a microservices function takes a long time to complete or if there is an intermittent problem with an upstream or downstream dependency. So, we need different tools to gather this type of information.

– We have an issue with auto-scaling metrics because they only look at individual microservices with a given set of attributes. So, they don’t give you a holistic view of the problem. For example, the application stack now exists in numerous locations and location types; we need a holistic viewpoint.

– A metric does not give this. For example, metrics track simplistic system states that indicate a service is running poorly or may be a leading indicator or an early warning signal. However, while those measures are easy to collect, they don’t turn out to be proper measures for triggering alerts.

Latency & Cloud Trace

Latency, in the context of applications, refers to the time it takes for a request to travel from the user to the server and back. Various factors, such as network delays, server processing time, and database queries influence it. Understanding latency is essential for developers to identify bottlenecks and optimize their applications for better performance.

Google Cloud Trace is a powerful tool provided by Google Cloud Platform that allows developers to analyze and diagnose application latency issues. By integrating Cloud Trace into their applications, developers can gain valuable insights into their code’s performance and identify areas for improvement.

Developers need to capture traces to analyze application latency effectively. Traces provide a detailed record of a request’s execution path, allowing developers to pinpoint the exact areas where latency occurs. With Cloud Trace, developers can easily capture and visualize traces in a user-friendly interface.

Auto-scaling metrics: Issues with dashboards

Useful only for a few metrics

So, these metrics are gathered and stored in time-series databases, and we have several dashboards to display these metrics. These dashboards were first built, and there weren’t many system metrics to worry about. You could have gotten away with 20 or so dashboards. But that was about it.

As a result, it was easy to see the critical data anyone should know about for any given service. Moreover, those systems were simple and did not have many moving parts. This contrasts with modern services that typically collect so many metrics that fitting them into the same dashboard is impossible.

Issues with aggregate metrics

So, we must find ways to fit all the metrics into a few dashboards. Here, the metrics are often pre-aggregated and averaged. However, the issue is that the aggregate values no longer provide meaningful visibility, even when we have filters and drill-downs. Therefore, we need to predeclare conditions that describe conditions we expect in the future. 

This is where we use instinctual practices based on past experiences and rely on gut feeling. Remember the network and software hero? It would help to avoid aggregation and averaging within the metrics store. On the other hand, we have percentiles that offer a richer view. Keep in mind, however, that they require raw data.

**Auto Scaling Observability: Any Question**

A: ) For auto-scaling observability, we take on an entirely different approach. They strive for other exploratory methods to find problems. Essentially, those operating observability systems don’t sit back and wait for an alert or something to happen. Instead, they are always actively looking and asking random questions to the observability system.

B: ) Observability tools should gather rich telemetry for every possible event, having full content of every request and then being able to store it and query it. In addition, these new auto-scaling observability tools are specifically designed to query against high-cardinality data. High cardinality allows you to interrogate your event data in any arbitrary way that we see fit. Now, we ask any questions about your system and inspect its corresponding state. 

**No predictions in advance**

C: ) Due to the nature of modern software systems, you want to understand any inner state and services without anticipating or predicting them in advance. For this, we need to gain valuable telemetry and use some new tools and technological capabilities to gather and interrogate this data once it has been collected. Telemetry needs to be constantly gathered in flexible ways to debug issues without predicting how failures may occur. 

D: ) Conditions affecting infrastructure health change infrequently and are relatively more straightforward to monitor. In addition, we have several well-established practices to predict, such as capacity planning and the ability to remediate automatically, e.g., auto-scaling in a Kubernetes environment. All of these can be used to tackle these types of known issues.

E: ) Due to its relatively predictable and slowly changing nature, the aggregated metrics approach monitors and alerts perfectly for infrastructure problems. So here, a metric-based system works well. Metrics-based systems and their associated signals help you see when capacity limits or known error conditions of underlying systems are being reached.

F: ) So, metrics-based systems work well for infrastructure problems that don’t change much but fall dramatically short in complex distributed systems. For these systems, you should opt for an observability and controllability platform. 

Kubernetes Autoscaling Architecture Builder

Build an autoscaling architecture and see which observability signals each component can use — and where its visibility stops.

1. Choose Your Architecture

Current model: HPA can use resource metrics such as CPU/memory and, when configured, custom or external metrics. Logs and traces provide investigation context but are not automatically HPA scaling signals.

2. Architecture Flow

APPLICATION
User requests, application workload and dependencies.
Pods Services
↓
TELEMETRY
Metrics + application metrics + logs
↓
OBSERVABILITY
Collection, correlation and analysis.
OpenTelemetry Prometheus Grafana
↓
AUTOSCALER
HPA evaluates configured metrics against targets.
↓
CAPACITY
Workload and infrastructure capacity changes.
Pod count Node count
↓
USER EXPERIENCE
Availability, latency, errors and application performance.
Engineering takeaway: Autoscaling and observability are related, but they are not the same thing. The autoscaler acts on configured signals; observability gives engineers the broader context needed to understand whether scaling actually improved the system. A healthy architecture therefore closes the loop: telemetry → decision → capacity → user experience → investigation.
ACI networks

ACI Networks

ACI Networks

Cisco Application Centric Infrastructure (ACI) is a software-defined networking architecture designed to connect applications, workloads, and network infrastructure through a common policy model. Rather than configuring every switch and interface independently, ACI allows network behaviour to be expressed through application requirements and translated into fabric-level configuration and policy.

At the heart of ACI is a leaf-spine fabric managed by the Application Policy Infrastructure Controller (APIC). Leaf switches provide connectivity to endpoints such as servers, virtual machines, containers, and external networks, while spine switches provide the high-speed fabric between leaf nodes. This architecture creates a predictable forwarding topology while centralising policy and operational control.

ACI’s application-centric model is built around concepts such as tenants, application profiles, endpoint groups (EPGs), contracts, and filters. Endpoint groups allow workloads with similar connectivity requirements to be grouped logically, independent of their physical location. Contracts then define the communication relationships permitted between those groups. This provides a policy-driven alternative to managing connectivity solely through individual VLANs, interfaces, and access-control entries.

Policy-based automation is one of ACI’s primary operational advantages. Administrators define the intended state of the network through APIC rather than manually configuring every underlying device. The controller translates those policies into the appropriate fabric configuration, helping maintain consistency and reducing repetitive configuration tasks across large environments.

ACI also provides a foundation for segmentation and security policy. Contracts can control communication between application tiers, while microsegmentation and endpoint-based policy can provide more granular controls where required. This allows security boundaries to follow application relationships rather than relying exclusively on physical network boundaries.

The ACI architecture is also designed for integration with virtualisation, containers, external networks, and cloud environments. Organisations can use ACI policy and fabric capabilities alongside technologies such as VMware, Kubernetes, and external Layer 2 or Layer 3 connectivity. In multi-environment architectures, this can help maintain consistent operational and security models as workloads move between different infrastructure domains.

ACI is particularly relevant in environments where application dependencies are complex and network changes occur frequently. In a traditional network, deploying a new application tier may require changes across VLANs, switch ports, routing, firewall policies, and load-balancing infrastructure. An application-centric model instead attempts to represent the required connectivity as a set of policies and relationships that can be applied consistently across the fabric.

Real-world deployments commonly include data centres, private cloud platforms, hybrid environments, and multi-tenant infrastructure. The value of ACI is not simply automation for its own sake; its greater objective is to connect network configuration with application intent, improving operational consistency, visibility, segmentation, and the speed at which infrastructure can respond to application requirements.

Understanding ACI therefore requires looking beyond individual Cisco switches and commands. The important concepts are the relationship between the physical fabric, APIC controller, endpoint groups, contracts, application profiles, and policy enforcement points. Together, these components form the foundation of Cisco’s application-centric approach to data-centre networking.
::

ACI Fabric Builder

Build the core Cisco ACI architecture and see how the fabric separates policy, forwarding, and application connectivity.

1. Select Fabric Components

ACI principle: APIC defines and distributes policy, while the fabric switches provide distributed forwarding. APIC is not in the data-plane path.

2. Fabric Architecture

Management / Policy
APIC Policy + Automation
↓
Fabric
Spine 1 Fabric Backbone
Spine 2 Fabric Backbone
Leaf 1 Endpoint Access
Leaf 2 Endpoint Access
↓
Workloads
Web Endpoint
App Endpoint
DB Endpoint
Engineering takeaway: ACI separates the policy model from the forwarding fabric. APIC provides centralized policy and management, while Leaf and Spine switches form the distributed forwarding infrastructure. VXLAN provides the overlay and the routed fabric underlay provides resilient IP connectivity and ECMP.

Highlights: ACI Networks

ACI Main Components

-APIC controllers and underlay network infrastructure are the main components of ACI. Due to specialized forwarding chips, hardware-based underlay switching in ACI has a significant advantage over software-only solutions.

-As a result of Cisco’s own ASIC development, ACI has many advanced features, including security policy enforcement, microsegmentation, dynamic policy-based redirection (allowing external L4-L7 service devices to be inserted into the data path), and detailed flow analytics—in addition to performance and flexibility.

-ACI underlays require Nexus 9000 switches exclusively. There are Nexus 9500 modular switches and Nexus 9300 fixed 1U to 2U models available. The spine function in ACI fabric is handled by certain models and line cards, while leaves can be handled by others, and some can even be used for both functions at the same time.The combination of different leaf switches within the same fabric is not limited.

ACI Networks

ACI Networks refers to the full architecture, fabric, and operational model behind Cisco Application Centric Infrastructure (ACI) — Cisco’s flagship SDN solution for modern data centers. ACI Networks combine policy‑driven automation, spine‑leaf switching, VXLAN overlays, centralized control, and application‑aware segmentation to create a programmable, scalable, and secure data‑center fabric.

This section continues your sequence perfectly (SD‑WAN → SASE → Security → Visibility → ACI).

What ACI Networks Are

ACI Networks are built on three pillars:

  • Policy Model — EPGs, contracts, tenants, and application profiles define intent.

  • Fabric Architecture — Spine‑leaf topology with VXLAN and distributed anycast gateways.

  • Centralized Control — APIC manages policies, fabric health, and automation.

ACI Networks replace manual configuration with intent‑based networking, where the network automatically aligns with application requirements.

ACI Network Architecture

An ACI Network consists of:

  • Leaf switches — where endpoints connect

  • Spine switches — high‑speed fabric backbone

  • APIC controllers — policy, automation, and fabric management

  • VXLAN overlays — scalable, multi‑tenant segmentation

  • Anycast gateways — distributed L3 gateways for optimal routing

This creates a high‑performance, low‑latency, horizontally scalable data‑center fabric.

ACI Network Components

  • Endpoint Groups (EPGs) — Logical groupings of workloads

  • Contracts — Define allowed communication between EPGs

  • Tenants — Multi‑domain isolation

  • Bridge Domains — L2 forwarding scopes

  • Application Profiles — Map app tiers to policies

  • Fabric Policies — Interface, routing, and access policies

These components form the foundation of ACI’s intent‑based model.

Security in ACI Networks

ACI Networks provide strong security through:

  • microsegmentation

  • EPG‑based isolation

  • contract‑based Zero Trust

  • identity‑aware endpoint classification

  • policy‑driven enforcement

Security becomes application‑centric, not IP‑centric.

ACI Network Visibility

Visibility is delivered through:

  • Nexus Dashboard Insights

  • Telemetry & health scores

  • Flow analytics

  • Endpoint tracking

This allows operators to detect drift, anomalies, misconfigurations, and performance issues across the fabric.

ACI Networks in Hybrid Cloud

ACI extends into cloud environments:

  • AWS

  • Azure

  • GCP

Using Cloud ACI, policies follow workloads across on‑prem and cloud, creating a unified segmentation and security model.

Outcome

ACI Networks deliver:

  • policy‑driven automation

  • scalable spine‑leaf architecture

  • secure segmentation

  • consistent hybrid‑cloud policies

  • deep visibility and assurance

  • application‑centric networking

ACI becomes the core SDN fabric for modern enterprise data centers.

Cisco Data Center Design

**The rise of virtualization**

Virtualization is creating a virtual — rather than actual — version of something, such as an operating system (OS), a server, a storage device, or network resources. Virtualization uses software that simulates hardware functionality to create a virtual system.

It is creating a virtual version of something like computer hardware. It was initially developed during the mainframe era. With virtualization, the virtual machine could exist on any host. As a result, Layer 2 had to be extended to every switch.

This was problematic for Larger networks as the core switch had to learn every MAC address for every flow that traversed it. To overcome this and take advantage of the convergence and stability of layer 3 networks, overlay networks became the choice for data center networking, along with introducing control plane technologies such as EVPM MPLS.

**The Cisco Data Center Design Transition**

The Cisco data center design has gone through several stages. First, we started with the Spanning Tree, moved to the Spanning Tree with vPCs, and then replaced the Spanning Tree with FabricPath. FabricPath is what is known as a MAC-in-MAC Encapsulation.

Today, in the data center, VXLAN is the de facto overlay protocol for data center networking. The Cisco ACI uses an enhanced version of VXLAN to implement both Layer 2 and Layer 3 forwarding with a unified control plane. Replacing SpanningTree with VXLAN, where we have a MAC-in-IP encapsulation, was a welcomed milestone for data center networking.

**Overlay networking with VXLAN**

VXLAN is an encapsulation protocol that provides data center connectivity using tunneling to stretch Layer 2 connections over an underlying Layer 3 network. VXLAN is the most commonly used protocol in data centers to create a virtual overlay solution that sits on top of the physical network, enabling virtual networks. The VXLAN protocol supports the virtualization of the data center network while addressing the needs of multi-tenant data centers by providing the necessary segmentation on a large scale.

Here, we are encapsulating traffic into a VXLAN header and forwarding between VXLAN tunnel endpoints, known as the VTEPs. With overlay networking, we have the overlay and the underlaying concept. By encapsulating the traffic into the overlay VXLAN, we now use the underlay, which in the ACI is provided by IS-IS, to provide the Layer 3 stability and redundant paths using Equal Cost Multipathing (ECMP) along with the fast convergence of routing protocols.

Example: Point to Point GRE

Cisco ACI Overview

Introduction to the ACI Networks

The base of the ACI network is the Cisco Application Centric Infrastructure Fabric (ACI)—the Cisco SDN solution for the data center. Cisco has taken a different approach from the centralized control plane SDN approach with other vendors and has created a scalable data center solution that can be extended to multiple on-premises, public, and private cloud locations.

The ACI networks have many components, including Cisco Nexus 9000 Series switches with the APIC Controller running in the spine leaf architecture ACI fabric mode. These components form the building blocks of the ACI, supporting a dynamic integrated physical and virtual infrastructure.

Enhanced Scalability and Flexibility:

One of the critical advantages of ACI networks is their ability to scale and adapt to changing business needs. Traditional networks often struggle to accommodate rapid growth or dynamic workloads, leading to performance bottlenecks. ACI networks, on the other hand, offer seamless scalability and flexibility, allowing businesses to quickly scale up or down as required without compromising performance or security.

Simplified Network Operations:

Gone are the days of manual network configurations and time-consuming troubleshooting. ACI networks introduce a centralized management approach, where policies and structures can be defined and automated across the entire network infrastructure. This simplifies network operations, reduces human errors, and enables IT teams to focus on strategic initiatives rather than mundane tasks.

Enhanced Security:

Network security is paramount in today’s threat landscape. ACI networks integrate security as a foundational element rather than an afterthought. With ACI’s microsegmentation capabilities, businesses can create granular security policies and isolate workloads, effectively containing potential threats and minimizing the impact of security breaches. This approach ensures that critical data and applications remain protected despite evolving cyber threats.

**Real-World Use Cases of ACI Networks**

  • Data Centers and Cloud Environments:

ACI networks have revolutionized data center and cloud environments, enabling businesses to achieve unprecedented agility and efficiency. By providing a unified management platform, ACI networks simplify data center operations, enhance workload mobility, and optimize resource utilization. Furthermore, ACI’s seamless integration with cloud platforms ensures consistent network policies and security across hybrid and multi-cloud environments.

  • Network Virtualization and Automation:

ACI networks are a game-changer for network virtualization and automation. By abstracting network functionality from physical hardware, ACI enables businesses to create virtual networks, provision services on-demand, and automate network operations. Streamlining network deployments accelerates service delivery, reduces costs, and improves overall performance.

Recap: Traditional Data Center 

Firstly, the Cisco data center design traditionally built our networks based on hierarchical data center topologies. This is often referred to as the traditional data center, which has a three-tier structure with an access layer, an aggregation layer, and a core layer. Historically, this design enabled substantial predictability because aggregation switch blocks simplified the spanning-tree topology. In addition, the need for scalability often pushed this design into modularity with ACI networks and ACI Cisco, which increased predictability.

Recap: The Challenges

However, although we increased predictability, the main challenge inherent in the three-tier models is that they were difficult to scale. As the number of endpoints increases and the need to move between segments increases, we need to span layer 2. This is a significant difference between the traditional and the ACI data centers.

 

ACI Networks

ACI Policy & EPG Lab

Test how ACI uses Endpoint Groups and Contracts to define application-centric communication instead of relying only on IP-based rules.

1. Define the Policy

2. Policy Decision

WEB-EPG Source
→ HTTPS
APP-EPG Destination
ALLOWED
WEB-EPG is permitted to communicate with APP-EPG using the HTTPS contract.
EPG Groups endpoints according to application or policy role.
Contract Defines the permitted communication relationship.
Tenant Provides policy and administrative isolation.
Microsegmentation Applies policy between logical application groups.
Engineering takeaway: In ACI, an EPG identifies a logical group of endpoints and a Contract defines the communication relationship between EPGs. The important shift is from thinking only in terms of IP addresses and VLANs to expressing application communication as policy.

The Journey to ACI

Our journey towards ACI started in the early 1990s when we examined the most traditional and well-known two—or three-layer network architecture. This Core/Aggregation/Access design was generally used and recommended for campus enterprise networks.

Layer 2 Connectivity:

At that time and in that environment, it delivered sufficient quality for typical client-server types of applications. The traditional design taken from campus networks was based on Layer 2 connectivity between all network parts, segmentation was implemented using VLANs, and the loop-free topology relied on the Spanning Tree Protocol (STP).

STP Limitations:

Scaling such an architecture implies growing broadcast and failure domains, which could be more beneficial for the resulting performance and stability. For instance, picture each STP Topology Change Notification (TCN) message causing MAC tables to age in the whole datacenter for a particular VLAN, followed by excessive BUM (Broadcast, Unknown Unicast, Multicast) traffic flooding until all MACs are relearned.

Spanning Tree Root Switch stp port states

**Designing around STP**

Before we delve into the Cisco ACI overview, let us first address some basics around STP design. The traditional Cisco data center design often leads to poor network design and human error. You don’t want a layer 2 segment between the data center unless you have the proper controls.

Although modularization is still desired in networks today, the general trend has been to move away from this design type, which evolves around a spanning tree, to a more flexible and scalable solution with VXLAN and other similar Layer 3 overlay technologies. In addition, the Layer 3 overlay technologies bring a lot of network agility, which is vital to business success.

VXLAN overlay

Agility refers to making changes, deploying services, and supporting the business at its desired speed. This means different things to different organizations. For example, a network team can be considered agile if it can deploy network services in a matter of weeks.

In others, it could mean that business units in a company should be able to get applications to production or scale core services on demand through automation with Ansible CLI or Ansible Tower.

Regardless of how you define agility, there is little disagreement with the idea that network agility is vital to business success. The problem is that network agility has traditionally been hard to achieve until now with the ACI data center. Let’s recap some of the leading Cisco data center design transitions to understand fully.

**Challenge: – Layer 2 to the Core**

The traditional SDN data center has gone through several transitions. Firstly, we had Layer 2 to the core. Then, from the access to the core, we had Layer 2 and not Layer 3. A design like this would, for example, trunk all VLANs to the core. For redundancy, you would manually prune VLANs from the different trunk links.

Our challenge with this approach of having Layer 2 to the core relies on the Spanning Tree Protocol. Therefore, redundant links are blocked. As a result, we don’t have the total bandwidth, leading to performance degradation and resource waste. Another challenge is to rely on topology changes to fix the topology.

Data Center Design

Data Center Stability

Layer 2 to the Core layer

STP blocks reduandant links

Manual pruning of VLANs

STP for topology changes

Efficient design

Spanning Tree Protocol does have timers to limit the convergence and can be tuned for better performance. Still, we rely on the convergence from the Spanning Tree Protocol to fix the topology, but the Spanning Tree Protocol was never meant to be a routing protocol.

Compared to other protocols operating higher up in the stack, they are designed to be more optimized to react to changes in the topology. However, STP is not an optimized control plane protocol, significantly hindering the traditional data center. You could relate this to how VLANs have transitioned to become a security feature. However, their purpose was originally for performance reasons.

**Required: – Routing to Access Layer**

To overcome these challenges and build stable data center networks, the Layer 3 boundary is pushed further to the network’s edge. Layer 3 networks can use the advances in routing protocols to handle failures and link redundancy much more efficiently.

It is a lot more efficient than Spanning Tree Protocol, which should never have been there in the first place. Then we had routing at the access. With this design, we can eliminate the Spanning Tree Protocol to the core and then run Equal Cost MultiPath (ECMP) from the access to the core.

We can run ECMP as we are now Layer 3 routing from the access to the core layer instead of running STP that blocks redundant links.  However, equal-cost multipath (ECMP) routes offer a simple way to share the network load by distributing traffic onto other paths.

ECMP is typically applied only to entire flows or sets of flows. Destination address, source address, transport level ports, and payload protocol may characterize a flow in this respect.

Data Center Design

Data Center Stability


Layer 3 to the Core layer

Routing protocol stability 

Automatic routing  convergence

STP for topology changes

Efficient design

**Key Point: – Equal Cost MultiPath (ECMP)**

Equal-cost Multipath (ECMP) has many advantages. First, ECMP gives us total bandwidth with equal-cost links. As we are routing, we no longer have to block redundant links to prevent loops at Layer 2. However, we still have Layer 2 in the network design and Layer 2 on the access layer; therefore, parts of the network will still rely on the Spanning Tree Protocol, which converges when there is a change in the topology.

So we may have Layer 3 from the access to the core, but we still have Layer 2 connections at the edge and rely on STP to block redundant links to prevent loops. Another potential drawback is that having smaller Layer 2 domains can limit where the application can reside in the data center network, which drives more of a need to transition from the traditional data center design.

The Layer 2 domain that the applications may use could be limited to a single server rack connected to one ToR or two ToR for redundancy with a Layer 2 interlink between the two ToR switches to pass the Layer 2 traffic.

These designs are not optimal, as you must specify where your applications are set, which limits agility. As a result, another critical Cisco data center design transition was the introduction of overlay data center designs.

**The Cisco ACI version**

Before Cisco ACI 4.1, the Cisco ACI fabric allowed only a two-tier (spine-and-leaf switch) topology. Each leaf switch is connected to every spine switch in the network, and there is no interconnection between leaf switches or spine switches.

Starting from Cisco ACI 4.1, the Cisco ACI fabric allows a multitier (three-tier) fabric and two tiers of leaf switches, which provides the capability for vertical expansion of the Cisco ACI fabric. This is useful for migrating a traditional three-tier architecture of core aggregation access that has been a standard design model for many enterprise networks and is still required today.

ACI fabric Details
Diagram: Cisco ACI fabric Details

The APIC Controller:

The ACI networks are driven by the Cisco Application Policy Infrastructure Controller ( APIC) database, which works in a cluster from the management perspective. The APIC is the centralized control point; you can configure everything in the APIC.

Consider the APIC to be the brains of the ACI fabric and server as the single source of truth for configuration within the fabric. The APIC controller is a policy engine and holds the defined policy, which tells the other elements in the ACI fabric what to do. This database allows you to manage the network as a single entity. 

In summary, the APIC is the infrastructure controller and is the main architectural component of the Cisco ACI solution. It is the unified point of automation and management for the Cisco ACI fabric, policy enforcement, and health monitoring. The APIC is not involved in data plane forwarding.

data center layout
Diagram: Data center layout: The Cisco APIC controller

The APIC represents the management plane, allowing the system to maintain the control and data plane in the network. The APIC is not the control plane device, nor does it sit in the data traffic path. Remember that the APIC controller can crash, and you still have forwarded in the fabric. The ACI solution is not an SDN centralized control plane approach. The ACI is a distributed fabric with independent control planes on all fabric switches. 

The Leaf and Spine 

Leaf-spine is a two-layer data center network topology for data centers that experience more east-west network traffic than north-south traffic. The topology comprises leaf switches (to which servers and storage connect) and spine switches (to which leaf switches connect).

In this two-tier Clos architecture, every lower-tier switch (leaf layer) is connected to each top-tier switch (Spine layer) in a full-mesh topology. The leaf layer consists of access switches connecting to devices like servers.

The Spine layer is the network’s backbone and interconnects all Leaf switches. Every Leaf switch connects to every spine switch in the fabric. The path is randomly chosen, so the traffic load is evenly distributed among the top-tier switches. Therefore, if one of the top-tier switches fails, it would only slightly degrade performance throughout the data center.

SDN data center
Diagram: Cisco ACI fabric checking.

Unlike the traditional Cisco data center design, the ACI data center operates with a Leaf and Spine architecture. Traffic now comes in through a device sent from an end host, known as a Leaf device.

We also have the Spine devices, which are Layer 3 routers with no unique hardware dependencies. In a primary Leaf and Spine fabric, every Leaf is connected to every Spine. Any endpoint in the fabric always has the same distance regarding hops and latency as every other internal endpoint.

The ACI Spine switches are Clos intermediary switches with many vital functions. Firstly, they exchange routing updates with leaf switches via Intermediate System-to-Intermediate System (IS-IS) and rapidly forward packets between them. They also provide endpoint lookup services to leaf switches through the Council of Oracle Protocol (COOP) and handle route reflection to the leaf switches using Multiprotocol BGP (MP-BGP).

Cisco ACI Overview
Diagram: Cisco ACI Overview.

The Leaf switches are the ingress/egress points for traffic into and out of the ACI fabric. They also provide end-host connectivity and are the connectivity points for the various endpoints that the Cisco ACI supports.

The spines act as a fast, non-blocking Layer 3 forwarding plane that supports Equal Cost Multipathing (ECMP) between any two endpoints in the fabric and uses overlay protocols such as VXLAN under the hood. VXLAN enables any workload to exist anywhere in the fabric, so we can now have workloads anywhere in the fabric without introducing too much complexity.

Required: ACI data center and ACI networks

This is a significant improvement to data center networking. We can now have physical or virtual workloads in the same logical Layer 2 domain, even running Layer 3 down to each ToR switch. The ACI data center is a scalable solution as the underlay is specifically built to be scalable as more links are added to the topology and resilient when links in the fabric are brought down due to, for example, maintenance or failure. 

ACI Networks: The Normalization event

VXLAN is an industry-standard protocol that extends Layer 2 segments over Layer 3 infrastructure to build Layer 2 overlay logical networks. The ACI infrastructure Layer 2 domains reside in the overlay, with isolated broadcast and failure bridge domains. This approach allows the data center network to grow without risking creating too large a failure domain. All traffic in the ACI fabric is normalized as VXLAN packets.

**Encapsulation Process**

ACI encapsulates external VLAN, VXLAN, and NVGRE packets in a VXLAN packet at the ingress. This is known as ACI encapsulation normalization. As a result, the forwarding in the ACI data center fabric is not limited to or constrained by the encapsulation type or overlay network. If necessary, the ACI bridge domain forwarding policy can be defined to provide standard VLAN behavior where required.

**Making traffic ACI-compatible**

As a final note in this Cisco ACI overview, let us address the normalization process. When traffic hits the Leaf, there is a normalization event. The normalization takes traffic from the servers to the ACI, making it ACI-compatible. Essentially, we are giving traffic sent from the servers a VXLAN ID to be sent across the ACI fabric.

Traffic is normalized, encapsulated with a VXLAN header, and routed across the ACI fabric to the destination Leaf, where the destination endpoint is. This is, in a nutshell, how the ACI Leaf and Spine work. We have a set of leaf switches that connect to the workloads and the spines that connect to the Leaf.

**VXLAN: Overlay Protocol** 

VXLAN is the overlay protocol that carries data traffic across the ACI data center fabric. A key point to this type of architecture is that the Layer 3 boundary is moved to the Leaf. This brings a lot of value and benefits to data center design. This boundary makes more sense as we must route and encapsulate this layer without going to the core layer.

ACI networks are revolutionizing how businesses connect and operate in the digital age. Focusing on application-centric infrastructure, they offer enhanced scalability, simplified network operations, and top-notch security. By leveraging ACI networks, businesses can unleash the full potential of their network infrastructure, ensuring seamless connectivity and staying ahead in today’s competitive landscape.

VXLAN & ACI Traffic Flow

Step through an ACI packet journey and separate the VXLAN overlay from the IP-based fabric underlay.

Packet Journey

1
Ingress Leaf The source endpoint sends traffic to its attached Leaf switch. The Leaf identifies the endpoint and applies the relevant policy.
2
VXLAN Encapsulation The Leaf encapsulates the original frame for transport across the ACI fabric overlay.
3
IP Underlay + ECMP The fabric underlay provides IP reachability between Leaves. Multiple Spine paths allow ECMP forwarding.
4
Egress Leaf The destination Leaf receives the encapsulated traffic and prepares it for delivery to the destination endpoint.
5
VXLAN Decapsulation The destination Leaf removes the fabric encapsulation and forwards the original traffic toward the endpoint.

Step 1 — Ingress Leaf

Packet View
Source Web Endpoint
Destination App Endpoint
Fabric State Original Frame
What is happening?
Traffic enters the ACI fabric through the source Leaf. The Leaf is the endpoint-facing point where classification and forwarding decisions begin.
Engineering takeaway: Think of ACI as two cooperating layers: the IP fabric underlay provides resilient transport between Leaf switches, while VXLAN provides the overlay used to carry tenant and endpoint traffic across that fabric. Spine switches provide fabric transit rather than acting as endpoint attachment points.
Cisco ACI

Service Level Objectives (SLOs): Customer-centric view

Service Level Objectives (SLOs)

Modern digital services are expected to be fast, available, and dependable—not simply operational. Service Level Objectives (SLOs) provide a measurable way to define what “reliable enough” means for users and to turn that expectation into an engineering target. Rather than treating availability or performance as abstract infrastructure numbers, effective SLOs connect technical behaviour directly to the experience a service is expected to deliver.

An SLO normally sits within a broader reliability model. A Service Level Indicator (SLI) is the measurement used to evaluate a service—for example, successful request rate, latency, availability, or data freshness. The SLO defines the target for that indicator, such as 99.9% successful requests over a rolling 30-day period. An SLA (Service Level Agreement) may then use selected service objectives as contractual commitments to customers. Keeping these concepts separate helps engineering teams distinguish between what they measure, what they aim to achieve, and what they formally promise.

Service Level Indicators (SLIs): Select measurements that represent the actual user experience. Useful SLIs can include request success rate, latency percentiles, availability, throughput, queue age, or freshness of data. The most valuable indicator is usually tied to a meaningful user journey rather than an isolated infrastructure component.

Service Level Objectives (SLOs): Define the desired level of reliability for each important SLI. For example, an API might target 99.9% successful requests and a 95th-percentile latency below a defined threshold over a rolling time window. SLOs should be measurable, meaningful, and realistic enough to guide engineering decisions.

Time Windows: Specify how reliability is evaluated. Rolling windows are commonly used because they provide a continuously updated view of recent service performance. The selected window should reflect the service’s usage patterns and the consequences of reliability failures.

Error Budgets: Translate an SLO into an explicit amount of acceptable unreliability. If a service has a 99.9% availability objective, the remaining 0.1% represents its error budget. This creates a practical mechanism for balancing reliability with feature development: when the budget is being consumed rapidly, engineering teams have evidence that reliability work should take priority.

Monitoring and Burn-Rate Alerting: Monitor SLO performance continuously and alert on meaningful consumption of the error budget rather than every individual metric fluctuation. Burn-rate approaches can identify situations where a service is consuming its available reliability budget significantly faster than expected, allowing teams to respond before the SLO is ultimately breached.

Capacity and Performance Planning: SLOs should influence architecture and capacity decisions. Increasing traffic, workload complexity, dependency failures, or resource saturation can all affect an SLO. Capacity planning therefore needs to consider not only whether infrastructure has enough resources, but whether the service can continue meeting its reliability and latency objectives as demand changes.

Continuous Improvement: Review SLOs regularly as the service, architecture, and user expectations evolve. An SLO that is too loose may fail to protect the user experience, while one that is unnecessarily strict can consume engineering capacity without providing meaningful additional value. SLOs should therefore be treated as living engineering agreements rather than permanent numbers.

The most effective SLO programmes move reliability discussions away from statements such as “the servers are healthy” and towards questions such as “Are users successfully completing the actions this service provides?” By combining meaningful SLIs, realistic SLOs, error budgets, observability, and burn-rate analysis, engineering teams can make reliability measurable and actionable.

Ultimately, SLOs are not simply another set of monitoring thresholds. They provide a framework for making informed trade-offs between reliability, performance, capacity, and development velocity—helping organisations build services that remain dependable as technology, traffic, and user expectations continue to evolve.

SLO Designer

Build a customer-focused Service Level Objective by selecting the service, reliability signal, target and measurement window.

1. Define the Objective

2. Resulting Reliability Objective

Service Level Objective
99.9%
The Public API should achieve 99.9% availability over a rolling 30-day measurement window.
SLI Availability
Target 99.9%
Window 30 days
User Perspective Successful service access
Engineering takeaway: An SLO should describe something users actually experience. The SLI is the measurement, the target defines the desired reliability, and the measurement window determines how that objective is evaluated over time.

Highlights: Service Level Objectives (SLOs)

Understanding Service Level Objectives

A ) Service Level Objectives, or SLOs, refer to predefined goals that outline the level of service a company aims to provide to its customers. They act as measurable targets that help organizations assess and improve service quality. To ensure consistent and reliable service delivery, SLOs define various metrics and performance indicators, such as response time, uptime, error rates, etc.

B ) Implementing SLOs offers numerous benefits for both businesses and customers. First, they provide a clear framework for service expectations, giving customers a transparent understanding of what to anticipate. Second, SLOs enable companies to align their internal goals with customer needs, fostering a customer-centric approach. Additionally, SLOs are crucial in driving accountability and continuous improvement within organizations, leading to enhanced operational efficiency and customer satisfaction.

C ) Setting Effective SLOs: To set effective SLOs, organizations need to consider several key factors. First, they must identify the critical metrics that directly impact customer experience. These could include response time, resolution time, or system availability. Second, SLOs should be realistic and achievable, taking into account the organization’s capabilities and resources. Moreover, SLOs should be regularly reviewed and adjusted to align with evolving customer expectations and business objectives.

D ) The Significance of SLOs in Business Success: Service Level Objectives are a theoretical and vital tool for achieving business success. Organizations can enhance customer satisfaction, loyalty, and retention by defining clear service goals.  This, in turn, leads to positive word-of-mouth, increased customer acquisition, and, ultimately, revenue growth. SLOs also enable businesses to identify and address service gaps proactively, mitigating potential issues before they escalate and impacting the overall customer experience.

Service Level Objectives (SLOs)

Service Level Objectives (SLOs) define the target reliability of a system — the level of performance you promise to users. They are measurable, time‑bound, and tied directly to user experience. SLOs are the backbone of modern reliability engineering, especially in cloud, microservices, Kubernetes, and auto‑scaling environments.

Below is a clean, structured section you can drop directly into your observability or SRE content.

What SLOs Are

An SLO is a numerical reliability target for a service. Examples:

  • “99.9% availability over 30 days”

  • “p95 latency under 200 ms”

  • “<0.1% error rate per rolling 7 days”

SLOs define how reliable the service must be to meet user expectations.

Core Components of SLOs

  • Service Level Indicators (SLIs) — The metrics you measure (latency, errors, availability).
  • Error Budget — The amount of failure allowed before reliability work must take priority.

  • Measurement Window — The time period (7 days, 30 days, 90 days).

  • User‑Centric Metrics — Metrics tied directly to user experience.

Together, these define the reliability contract between engineering and the business.

Examples of Common SLOs

Availability SLOs

  • 99.9% service uptime

  • 99.99% API availability

  • 99.5% successful login rate

Latency SLOs

  • p95 < 150 ms

  • p99 < 300 ms

  • request processing < 50 ms for 90% of calls

Error Rate SLOs

  • <0.1% failed requests

  • <0.5% 5xx responses

  • <1% authentication failures

Throughput SLOs

  • maintain 10k requests/sec with <1% degradation

Why SLOs Matter

SLOs provide:

  • predictable reliability

  • clear engineering priorities

  • alignment between product and operations

  • data‑driven decisions

  • a safety margin via error budgets

  • a way to balance innovation vs stability

They prevent teams from over‑engineering reliability or under‑investing in stability.

SLOs in Observability

SLOs integrate directly with observability systems:

  • metrics track SLI performance

  • logs show failure causes

  • traces reveal latency bottlenecks

  • dashboards visualize SLO burn rate

  • alerts trigger when error budgets burn too fast

This creates a closed feedback loop between reliability and monitoring.

SLOs in Auto‑Scaling & Cloud Systems

SLOs guide scaling decisions:

  • scale up when latency SLO is threatened

  • scale down when error rate stabilizes

  • pre‑scale before peak traffic

  • adjust HPA/VPA thresholds based on SLO burn rate

SLOs ensure auto‑scaling protects user experience, not just CPU metrics.

SLO Best Practices

  • define SLOs based on real user expectations

  • avoid unrealistic targets (e.g., 100% uptime)

  • use rolling windows (30 days is standard)

  • track SLO burn rate daily

  • tie error budgets to release velocity

  • automate SLO dashboards and alerts

SLOs should be simple, measurable, and meaningful.

Outcome

Implementing SLOs delivers:

  • predictable reliability

  • better user experience

  • controlled release velocity

  • intelligent auto‑scaling

  • reduced firefighting

  • clear operational priorities

  • measurable service health

SLOs become the north star of modern reliability engineering.

**Service Level Objectives**

Key Note: We need to start thinking differently about things than we have in the past to make sure our services are reliable with complex systems. It might be your responsibility to maintain a globally distributed service with thousands of moving parts, or it might just be to keep a few virtual machines. No matter how far removed humans are from those things, they almost certainly rely on them at some point. It is also necessary to consider things from the perspective of human users once you have considered their needs.

SLOs are percentages you use to help drive decision-making, while SLAs are promises to customers that include compensation in case you do not meet your targets. Violations of your SLO generate data you use to evaluate the reliability of your service. Whenever you violate an SLO, you can choose to take action.

I reiterate that SLOs are not contracts; they are objectives. You are free to update or change your targets at any time. In the world, things will change, and your service’s operations may also change.

**The Mechanics of Managed Instance Groups**

At its core, a managed instance group is a collection of VM instances that are treated as a single entity. This allows for seamless scaling, load balancing, and automated updates. By utilizing Google Cloud’s MIGs, you can configure these groups to dynamically adjust the number of instances in response to changes in demand, ensuring optimal performance and cost-efficiency. This flexibility is crucial for businesses that experience fluctuating workloads, as it allows for resources to be used judiciously, meeting service level objectives without unnecessary expenditure.

—

**Achieving Service Level Objectives with MIGs**

Service level objectives (SLOs) are critical benchmarks that dictate the expected performance and availability of your applications. Managed instance groups help you meet these SLOs by offering auto-healing capabilities, ensuring that any unhealthy instances are automatically replaced. This minimizes downtime and maintains the reliability of your services. Additionally, Google Cloud’s integration with load balancing allows for efficient distribution of traffic, further enhancing your application’s performance and adherence to SLOs.

—

**Advanced Features for Enhanced Management**

Google Cloud’s managed instance groups come equipped with advanced features that elevate their utility. With instance templates, you can standardize the configuration of your VM instances, ensuring consistency across your deployments. The use of regional managed instance groups provides additional resilience, as instances can be spread across multiple zones, safeguarding against potential outages in a single zone. These features collectively empower you to build robust, fault-tolerant applications tailored to your specific needs.

Managed Instance Group

Why are SLOs Important?

SLOs play a vital role in ensuring customer satisfaction and meeting business objectives. Here are a few reasons why SLOs are essential:

1. Accountability: SLOs provide a framework for holding service providers accountable for meeting the promised service levels. They establish a baseline for evaluating the performance and quality of the service.

2. Customer Experience: By setting SLOs, businesses can align their service offerings with customer expectations. This helps deliver a superior customer experience, foster customer loyalty, and gain a competitive edge in the market.

3. Performance Monitoring and Improvement: SLOs enable businesses to monitor their services’ performance and continuously identify improvement areas. Regularly tracking SLO metrics allows for proactive measures and optimizations to enhance service reliability and availability.

Critical Elements of SLOs:

To effectively implement SLOs, it is essential to consider the following key elements:

1. Metrics: SLOs should be based on relevant, measurable metrics that accurately reflect the desired service performance. Standard metrics include response time, uptime percentage, error rate, and throughput.

2. Targets: SLOs must define specific targets for each metric, considering customer expectations, industry standards, and business requirements. Targets should be achievable yet challenging enough to drive continuous improvement.

3. Monitoring and Alerting: Establishing robust monitoring and alerting mechanisms allows businesses to track the performance of their services in real time. This enables timely intervention and remediation in case of deviations from the defined SLOs.

4. Communication: Effective communication with customers is crucial to ensure transparency and manage expectations. Businesses should communicate SLOs, including the metrics, targets, and potential limitations, to foster trust and maintain a healthy customer-provider relationship.

**The Value of SRE Teams**

Site Reliability Engineering (SRE) teams have tools such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets that can guide them on the road to building a reliable system with the customer viewpoint as the metric. These new tools or technologies form the basis for reliability in distributed system and are the core building blocks of a reliable stack that assist with baseline engineering. The first thing you need to understand is the service’s expectations. This introduces the areas of service-level management and its components.

**The Role of Service-Level Management**

The core concepts of service level management are Service Level Agreement (SLA), Service Level Objectives (SLO), and Service Level Indicators (SLIs). The common indicators used are Availability, latency, duration, and efficiency. Monitoring these indicators to catch problems before your SLO is violated is critical. These are the cornerstone of developing a good SRE practice.

    • SLI: Service level Indicator: A well-defined measure of “successful enough.” It is a quantifiable measurement of whether a given user interaction was good enough. Did it meet the users’ expectations? Does a web page load within a specific time? This allows you to categorize whether a given interaction is good or bad.
    • SLO: Service level objective: A top-line target for a fraction of successful interactions.
    • SLA: Service level agreement: consequences. It’s more of a legal construct. 
Interactive SRE Lab 02

Error Budget & Burn Lab

See how the same SLO can move from healthy to exhausted as service failures accumulate. The key concept is burn rate: how quickly the available error budget is being consumed.

Error budget
43.2 min
Burn rate
2×
Budget consumed
20%
Error-budget consumption
0% 50% 100%
Budget under pressure Reliability work should be investigated before the budget is exhausted.
Engineering takeaway: An error budget turns reliability into an operational decision. A high burn rate means the service is consuming its permitted failure budget faster than planned. In practice, teams can use burn-rate alerts to trigger investigation, slow risky releases, or prioritize reliability work.

Service Level Objectives (SLOs)

Site Reliability Engineering (SRE)

Google pioneered SRE to create more scalable and reliable large-scale systems. SRE has become one of today’s most valuable software innovation opportunities. It is a concrete, opinionated implementation of the DevOps philosophy. The main goal is to create scalable and highly reliable software systems.

According to Benjamin Treynor Sloss, the founder of Google’s Site Reliability Team, “SRE is what happens when a software engineer is tasked with what used to be called operations.”

So, reliability is not so much a feature as a practice that must be prioritized and considered from the very beginning. It should not be added later, for example, when a system or service is in production. Reliability is the essential feature of any system, and it’s not a feature that a vendor can sell you.

Personal Note:

So, if someone tries to sell you an add-on solution called Reliability, don’t buy it, especially if they offer 100% reliability. Nothing can be 100% reliable all the time. If you strive for 100% reliability, you will miss out on opportunities to perform innovative tasks and the need to experiment and take risks that can help you build better products and services. 

Nothing can be 100% reliable all the time

Components of a Reliable System

### Distributed systems

At its core, a distributed system is a network of independent computers that work together to achieve a common goal. These systems can be spread across multiple locations and connected through communication networks. The architecture of distributed systems can vary widely, ranging from client-server models to peer-to-peer networks. One of the key features of distributed systems is their ability to provide redundancy and fault tolerance, ensuring that if one component fails, the system as a whole continues to function.

### Building Reliable Systems

To build reliable systems that can tolerate various failures, the system needs to be distributed so that a problem in one location doesn’t mean your entire service stops operating. So you need to build a system that can handle, for example, a node dying or perform adequately with a particular load.

To create a reliable system, you need to understand it fully and what happens when the different components that make up the system reach certain thresholds. This is where practices such as Chaos engineering kubernetes can help you.

### Chaos Engineering 

We can have practices like Chaos Engineering that can confirm your expectations, give you confidence in your system at different levels, and prove you can have certain tolerance levels to Reliability. Chaos Engineering allows you to find weaknesses and vulnerabilities in complex systems. It is an important task that can be automated into your CI/CD pipelines.

You can have various Chaos Engineering verifications before you reach production. These tests, such as load and Latency tests, can all be automated with little or no human interaction. Site Reliability Engineering (SRE) teams often use Chaos Engineering to improve resilience, which must be part of your software development/deployment process.  

### Integrating Chaos Engineering with Service Mesh

Integrating chaos engineering with a service mesh brings numerous benefits to organizations striving for resilience and reliability. Firstly, it enhances fault tolerance by exposing vulnerabilities before they become critical issues. Secondly, it provides a deeper understanding of system behavior under duress, enabling teams to optimize service performance and reliability. Lastly, it fosters a culture of experimentation and learning, encouraging teams to continuously improve and innovate.

 Perception: Customer-Centric View

Reliability is all about perception. Suppose the user considers your service unreliable. In that case, you will lose consumer trust because of poor service perception, so it’s important to provide consistency with your services as much as possible. For example, it’s OK to have some outages. Outages are expected, but you can’t have them all the time and for long durations.

Users expect to have outages at some point in time, but not for so long. User Perception is everything; if the user thinks you are unreliable, you are. Therefore, you need to have a customer-centric view, and using customer satisfaction is a critical metric to measure.

This is where the critical components of service management, such as Service Level Objectives (SLO) and Service Level Indicators (SLI), come into play. It would be best if you found a balance between Velocity and Stability. You can’t stop innovation, but you can’t take too many risks. An Error Budget will help you with Site Reliability Engineering (SRE) principles. 

Users experience Static thresholds.

User experience means different things to different groups of users. We now have a model where different service users may be routed through the system in other ways, using various components and providing experiences that can vary widely. We also know that the services no longer tend to break in the same few predictable ways over and over.

With complex microservices and many software interactions, we have many unpredictable failures that we have never seen before. These are often referred to as black holes. We should have a few alerts triggered by only focusing on symptoms that directly impact user experience and not because a threshold was reached.

Example: Issues with Static Thresholds

1.If your POD network reaches a certain threshold, this does not tell you anything about user experience. You can’t rely on static thresholds anymore, as they have no relationship to customer satisfaction.

2.If you use static thresholds, they can’t reliably indicate any issues with user experience. Alerts should be set up to detect failures that impact user experience. Traditional monitoring falls short of trying this as it usually has predefined dashboards that look for something that has happened before.

3.This brings us back to the challenges with traditional metrics-based monitoring; we rely on static thresholds to define optimal system conditions, which have nothing to do with user experience. However, modern systems change shape dynamically under different workloads. Static thresholds for monitoring can’t reflect impacts on user experience. They lack context and are too coarse.

How to Approach Reliability 

**New tools and technologies**

– 1: We have new tools, such as distributed tracing. What is the best way to find the bottleneck if the system becomes slow? Here, you can use Distributed Tracing and Open Telemetry. Tracing helps us instrument our system so we figure out where the time has been spent. It can be used across distributed microservice architecture to troubleshoot problems. Open Telemetry provides a standardized way of instrumenting our system and providing those traces.

– 2: have already touched on Service Level Objectives, Indicators, and Error Budget. You want to know why and how something has happened. So we don’t just want to know when something has happened and then react to an event that is not looking from the customer’s perspective.

– 3: need to understand if we are meeting the Service Level Agreement (SLA) by gathering the number and frequency of the outages and any performance issues. Service Level Objectives (SLO) and Service Level Indicators (SLI) can assist you with measurements. 

– 4: Level Objectives (SLO) and Service Level Indicators (SLI) assist you with measurements. They also offer a tool for better system reliability and form the base for the Reliability Stack. SLIs and SLOs help us interact with Reliability differently and provide a path for building a reliable system.

So now we have the tools and a disciple to use the tools within. Can you recall what that disciple is? The discipline is Site Reliability Engineering (SRE)

Example: Distributed Tracing with Cloud Trace

**SLO-based approach to reliability**

If you’re too reliable all the time, you’re also missing out on some of the fundamental features that SLO-based approaches give you. The main area you will miss is the freedom to do what you want, test, and innovate. If you’re too reliable, you’re missing out on opportunities to experiment, perform chaos engineering, ship features quicker than before, or even introduce structured downtime to see how your dependencies react.

To learn a system, you need to break it. So, if you are 100% reliable, you can’t touch your system, so you will never truly learn and understand your system. You want to give your users a good experience, but you’ll run out of resources in various ways if you try to ensure this good experience happens 100% of the time. SLOs let you pick a target that lives between those two worlds.

**Balance velocity and stability**

You can’t just have Reliability; you must also have new features and innovation. Therefore, you need to find a balance between velocity and stability. We need to balance Reliability with other features you have and are proposing to offer. Suppose you have access to a system with a fantastic feature that doesn’t work. The users who have the choice will leave.

Site Reliability Engineering is the framework for balancing velocity and stability. How do you know what level of Reliability you need to provide your customer? This all goes back to the business needs that reflect the customer’s expectations. With SRE, we have a customer-centric approach.

The primary source of outages is making changes even when the changes are planned. This can come in many forms, such as pushing new features, applying security patches, deploying new hardware, and scaling up to meet customer demand, which will significantly impact if you strive for a 100% reliability target. 

There will always be changes

If nothing changes to the physical/logical infrastructure or other components, we will not have bugs. We can freeze our current user base and never have to scale the system. In reality, this will not happen. There will always be changes. So it would be best if you found a balance.

Service Level Objectives (SLOs) are a cornerstone for delivering reliable and high-quality services in today’s technology-driven world. By setting measurable targets, businesses can align their service performance with customer expectations, drive continuous improvement, and ultimately enhance customer satisfaction. 

Implementing and monitoring SLOs allows companies to proactively address issues, optimize service delivery, and stay ahead of the competition. By embracing SLOs, companies can pave the way for successful service delivery and long-term growth.

Interactive SRE Lab 03

Customer-Centric Reliability Investigation

A service can look healthy from infrastructure metrics while customers are still experiencing failures. Investigate the difference between system health and user experience.

Infrastructure View

CPU utilisation 58%
Memory 62%
Availability 99.98%
Server errors 0.02%

Customer Experience

p95 response time 2.8s
Successful requests 98.7%
Checkout experience Degraded
User impact High
Infrastructure looks healthy — users do not

CPU and memory are within normal ranges, but elevated latency is degrading the customer experience. A CPU dashboard alone would miss the SLO risk.

auto scaling observability

Observability vs Monitoring

Observability vs Monitoring

Modern distributed systems generate enormous amounts of telemetry across applications, infrastructure, networks, databases, containers, and cloud services. As systems become more dynamic, simply knowing that something is “up” is no longer enough. Engineers need to understand what is happening, where it is happening, why it is happening, and which users or services are affected. This is where monitoring and observability play complementary roles.

Although the terms are often used interchangeably, they describe different capabilities. Monitoring is primarily concerned with collecting known signals and detecting known conditions—for example, high CPU utilisation, elevated error rates, or a service becoming unavailable. Observability provides the telemetry, context, and exploration capabilities required to investigate system behaviour, including problems that were not fully anticipated when the monitoring rules were created.

What is Observability?

Observability is the ability to understand the internal state and behaviour of a system by examining the telemetry it produces. Modern observability commonly combines metrics, logs, traces, and events, with additional context such as service metadata, deployment information, infrastructure state, and correlation identifiers. This allows engineers to move from a symptom to the underlying cause across complex service dependencies.

The Role of Monitoring:

Monitoring continuously measures selected signals and compares them with expected conditions, thresholds, or objectives. Examples include request error rate, latency, CPU utilisation, memory pressure, disk capacity, queue depth, or service availability. Monitoring is particularly effective for detecting known failure modes and triggering alerts when operational conditions require attention.

Known Conditions vs Unknown Questions:

Monitoring works well when engineers already know what constitutes an important failure condition. An alert might fire when API errors exceed a defined percentage or latency crosses an established threshold. Observability extends the investigation by allowing engineers to explore telemetry and ask questions that were not necessarily encoded into an individual alert—for example, whether failures are isolated to one service version, geographic region, customer segment, or downstream dependency.

Telemetry and Context:

The value of observability comes from combining signals rather than viewing them in isolation. A metric may show that request latency has increased, logs may reveal application errors, and distributed traces may identify a slow downstream database call. Correlation between these signals provides the context required to understand the relationship between symptoms and causes.

Detection and Investigation:

Monitoring is often the first mechanism that tells an engineering team that something requires attention. Observability then supports the investigation. A typical workflow might begin with an alert for elevated HTTP 5xx responses, move to service-level metrics, examine logs for the affected instances, and use distributed tracing to follow failed requests across multiple services.

Observability and Monitoring Work Together:

These approaches should not be treated as competing technologies. Effective operations normally require both. Monitoring provides visibility into known operational conditions and reliable alerting, while observability provides the depth and context needed to investigate complex or previously unknown behaviour. Monitoring can tell you that a service is unhealthy; observability helps you understand why.

For modern cloud-native and distributed environments, the goal is therefore not to choose between monitoring and observability. A mature platform combines meaningful monitoring and alerting with rich, correlated telemetry, enabling engineers to detect important problems quickly and investigate them efficiently. Together, they form an essential foundation for reliability engineering, incident response, performance optimisation, and continuous improvement.

Interactive Observability Lab 01

Monitoring vs Observability Decision Lab

Monitoring is excellent at detecting known failure conditions. Observability helps engineers investigate why a system behaves the way it does — including problems that were not predicted in advance.

Scenario 1 of 3
Scenario 1

CPU utilisation has exceeded 85% for five minutes.

The operations team wants an immediate alert whenever the threshold is exceeded.

0 / 3

Complete the scenarios to see your result.

Highlights: Observability vs Monitoring

Observability vs Monitoring

Monitoring and Distributed Systems

By utilizing distributed architectures, the cloud native ecosystem allows organizations to build scalable, resilient, and novel software architectures. However, the ever-changing nature of distributed systems means that previous approaches to monitoring can no longer keep up. The introduction of containers made the cloud flexible and empowered distributed systems.

Nevertheless, the ever-changing nature of these systems can cause them to fail in many ways. Distributed systems are inherently complex, and, as systems theorist Richard Cook notes, “Complex systems are intrinsically hazardous systems.”

Cloud-native systems require a new approach to monitoring, one that is open-source compatible, scalable, reliable, and able to control massive data growth. However, cloud-native monitoring can’t exist in a vacuum; it must be part of a broader observability strategy.

**Gaining Observability**

Key Features of Observability:

1. High-dimensional data collection: Observability involves collecting a wide variety of data from different system layers, including metrics, logs, traces, and events. This comprehensive data collection provides a holistic view of the system’s behavior.

2. Distributed tracing: Observability allows tracing requests as they flow through a distributed system, enabling engineers to understand the path and identify performance bottlenecks or errors.

3. Contextual understanding: Observability emphasizes capturing contextual information alongside the data, enabling teams to correlate events and understand the impact of changes or incidents.

Benefits of Observability:

1. Faster troubleshooting: By providing detailed insights into system behavior, observability helps teams quickly identify and resolve issues, minimizing downtime and improving system reliability.

2. Proactive monitoring: Observability allows teams to detect potential problems before they become critical, enabling proactive measures to prevent service disruptions.

3. Improved collaboration: With observability, different teams, such as developers, operations, and support, can have a shared understanding of the system’s behavior, leading to improved collaboration and faster incident response.

**Gaining Monitoring**

On the other hand, monitoring focuses on collecting and analyzing metrics to assess a system’s health and performance. It involves setting up predefined thresholds or rules and generating alerts based on specific conditions.

Key Features of Monitoring:

1. Metric-driven analysis: Monitoring relies on predefined metrics collected and analyzed to measure system performance, such as CPU usage, memory consumption, response time, or error rates.

2. Alerting and notifications: Monitoring systems generate alerts and notifications when predefined thresholds or rules are violated, enabling teams to take immediate action.

3. Historical analysis: Monitoring systems provide historical data, allowing teams to analyze trends, identify patterns, and make informed decisions based on past performance.

Benefits of Monitoring:

1. Performance optimization: Monitoring helps identify performance bottlenecks and inefficiencies within a system, enabling teams to optimize resources and improve overall system performance.

2. Capacity planning: By monitoring resource utilization and workload patterns, teams can accurately plan for future growth and ensure sufficient resources are available to meet demand.

3. Compliance and SLA enforcement: Monitoring systems help organizations meet compliance requirements and enforce service level agreements (SLAs) by tracking and reporting on key metrics.

Observability & Monitoring: A Unified Approach

While observability and monitoring differ in their approaches and focus, they are not mutually exclusive. When used together, they complement each other and provide a more comprehensive understanding of system behavior.

Observability enables teams to gain deep insights into system behavior, understand complex interactions, and troubleshoot issues effectively. Conversely, monitoring provides a systematic approach to tracking predefined metrics, generating alerts, and ensuring the system meets performance requirements.

Combining observability and monitoring can help organizations create a robust system monitoring and management strategy. This integrated approach empowers teams to quickly detect, diagnose, and resolve issues, improving system reliability, performance, and customer satisfaction.

Application Latency & Cloud Trace

A: – Latency, in simple terms, refers to the delay between sending a request and receiving a response. It can be caused by various factors, such as network congestion, server processing time, or inefficient code execution. Understanding the different components contributing to latency is essential for optimizing application performance.

B: – Google Cloud Trace is a powerful diagnostic tool provided by Google Cloud Platform. It allows developers to visualize and analyze latency data for their applications. By instrumenting code and capturing trace data, developers gain valuable insights into the performance bottlenecks and can take proactive measures to improve latency.

C: – To start capturing traces in your application, you need to integrate the Cloud Trace API into your codebase. Once integrated, Cloud Trace collects detailed latency data, including information about the various services and resources used to process a request. This data can then be visualized and analyzed through the user-friendly Cloud Trace interface.

The Starting Point: Observability vs Monitoring

You need to measure and gather the correct event information in your environment, which will be done with several tools. This will let you know what is affecting your application performance and infrastructure. As a good starting point, there are four golden signals for Latency, saturation, traffic, and errors. These are Google’s Four Golden Signals. The four most important metrics to keep track of are: 

      1. Latency: How long it takes to serve a request
      2. Traffic: The number of requests being made.
      3. Errors: The rate of failing requests. 
      4. Saturation: How utilized the service is.

So now we have some guidance on what to monitor and let us apply this to Kubernetes to, for example, let’s say, a frontend web service that is part of a tiered application, we would be looking at the following:

      1. How many requests is the front end processing at a particular point in time,
      2. How many 500 errors are users of the service received, and 
      3. Does the request overutilize the service?

We already know that monitoring is a form of evaluation that helps identify the most practical and efficient use of resources. With monitoring, we observe and check the progress or quality of something over time. Within this, we have metrics, logs, and alerts. Each has a different role and purpose.

**Monitoring: The role of metrics**

Metrics are related to some entity and allow you to view how many resources you consume. Metric data consists of numeric values instead of unstructured text, such as documents and web pages. Metric data is typically also a time series, where values or measures are recorded over some time. 

Available bandwidth and latency are examples of such metrics. Understanding baseline values is essential. Without a baseline, you will not know if something is happening outside the norm.

Note: Average Baselines

What are the average baseline values for bandwidth and latency metrics? Are there any fluctuations in these metrics? How do these values rise and fall during normal operations and peak usage? This may change over different days, weeks, and months.

If you notice a rise in these values during normal operations, this would be deemed abnormal and should act as a trigger that something could be wrong and needs to be investigated. Remember that these values should not be gathered as a once-off but can be gathered over time to understand your application and its underlying infrastructure better.

**Monitoring: The role of logs**

Logging is an essential part of troubleshooting application and infrastructure performance. Logs give you additional information about events, which is important for troubleshooting or discovering the root cause of the events. Logs will have much more detail than metrics, so you will need some way to parse the logs or use a log shipper.

A typical log shipper will take these logs from the standard out in a Docker container and ship them to a backend for processing.

Note: Example Log Shipper

FluentD or Logstash has pros and cons. The group can use it here and send it to a backend database, which could be the ELK stack ( Elastic Search). Using this approach, you can add different things to logs before sending them to the backend. For example, you can add GEO IP information. This will add richer information to the logs that can help you troubleshoot.

Understanding VPC Flow Logs

VPC Flow Logs is a feature provided by Google Cloud that captures network traffic metadata within a Virtual Private Cloud (VPC) network. This metadata includes source and destination IP addresses, protocol, port, and more. By enabling VPC Flow Logs, administrators can gain visibility into the network traffic patterns and better understand the communication flow within their infrastructure.

We can leverage data visualization tools to make the analysis more visually appealing and easier to comprehend. Google Cloud provides various options for creating interactive and informative dashboards, such as Data Studio and Cloud Datalab. These dashboards can display network traffic trends, highlight critical metrics, and aid in identifying patterns or anomalies that might require further investigation.

**Monitoring: The role of alerting**

Then we have the alerting, and it would be best to balance how you monitor and what you alert on. So, we know that alerting is not always perfect, and getting the right alerting strategy in place will take time. It’s not a simple day-one installation and requires much effort and cross-team collaboration.

You know that alerting on too much can cause alert fatigue. We are all too familiar with the problems alert fatigue can bring and the tensions it can create in departments.

To minimize this, consider Service Level Objectives (SLOs) for alerts. SLOs are measurable characteristics such as availability, throughput, frequency, and response times. They are the foundation for a reliability stack. Also, it would be best if you considered alert thresholds. If these are too short, you will get a lot of false positives on your alerts. 

Monitoring is not enough.

Even with all of these in place, monitoring is not enough. Due to the sheer complexity of today’s landscape, you need to consider and think differently about the tools you use and how you use the intelligence and data you receive from them to resolve issues before they become incidents.  That monitoring by itself is not enough.

The tool used to monitor is just a tool that probably does not cross technical domains, and different groups of users will administer each tool without a holistic view. The tools alone can take you only half the way through the journey.  Also, what needs to be addressed is the culture and the traditional way of working in silos. A siloed environment can affect the monitoring strategy you want to implement. Here, you can look at an observability platform.

The Foundation of GKE-Native Monitoring

GKE-Native Monitoring builds upon the robust foundation of Prometheus and Stackdriver, providing a seamless integration that simplifies observability within GKE clusters. By harnessing the strengths of these industry-leading monitoring solutions, GKE-Native Monitoring offers a robust and comprehensive monitoring experience.

Under the umbrella of GKE-Native Monitoring, users gain access to a rich set of features designed to enable fine-grained visibility and control. These include customizable dashboards, real-time metrics, alerts, and horizontal pod autoscaling. With these tools, developers and operators can easily monitor the health, performance, and resource utilization of their GKE clusters.

Observability vs Monitoring

When it comes to observability vs. monitoring, we know that monitoring can detect problems and tell you if a system is down, and when your system is UP, Monitoring doesn’t care. Monitoring only cares when there is a problem. The problem has to happen before monitoring takes action. It’s very reactive. So, if everything is working, monitoring doesn’t care.

On the other hand, we have an observability platform, which is a more proactive practice. It’s about what and how your system and services are doing. Observability lets you improve your insight into how complex systems work and quickly get to the root cause of any problem, known or unknown.

Observability is best suited for interrogating systems to explicitly discover the source of any problem, along any dimension or combination of dimensions, without first predicting. This is a proactive approach.

## Pillars of Observability

This is achieved by combining logs, metrics, and traces. So, we need data collection, storage, and analysis across these domains while also being able to perform alerting on what matters most. Let’s say you want to draw correlations between units like TCP/IP packets and HTTP errors experienced by your app.

The Observability platform pulls context from different sources of information, such as logs, metrics, events, and traces, into one central context. Distributed tracing adds a lot of value here.

Also, when everything is placed into one context, you can quickly switch between the necessary views to troubleshoot the root cause. Viewing these telemetry sources with one single pane of glass is an excellent key component of any observability system. 

## Known and Unknown vs Unknown and Unknown 

Monitoring automatically reports whether known failure conditions are occurring or are about to occur. In other words, it is optimized for reporting on unknown conditions about known failure modes, which are referred to as known unknowns. In contrast, Observability is centered around discovering if and why previously unknown failure modes may be occurring, in other words, to find unknown unknowns.

The monitoring-based approach of metrics and dashboards is an investigative practice that relies on humans’ experience and intuition to detect and understand system issues. This is okay for a simple legacy system that fails in predictable ways, but the instinctual technique falls short for modern systems that fail in unpredictable ways.

With modern applications, the complexity and scale of their underlying systems quickly make that approach unattainable, and we can’t rely on hunches. Observability tools differ from traditional monitoring tools because they enable engineers to investigate any system, no matter how complex. You don’t need to react to a hunch or have intimate system knowledge to generate a hunch.

Monitoring vs Observability: Working together?

Monitoring helps engineers understand infrastructure concerns, while observability helps engineers understand software concerns. So, Observability and Monitoring can work together. First, the infrastructure does not change too often, and when it fails, it will fail more predictably. So, we can use monitoring here.

This is compared to software system states that change daily and are unpredictable. Observability fits this purpose. The conditions that affect infrastructure health change infrequently and are relatively more straightforward to predict. 

We have several well-established practices to expect, such as capacity planning and the ability to remediate automatically (e.g., auto-scaling in a Kubernetes environment). All of these can be used to tackle these types of known issues. 

Monitoring and infrastructure problems

Due to its relatively predictable and slowly changing nature, the aggregated metrics approach monitors and alerts perfectly for infrastructure problems. So here, a metric-based system works well. Metrics-based systems and their associated alerts help you see when capacity limits or known error conditions of underlying systems are being reached.

Now, we need to look at monitoring the Software and have access to high-cardinality fields. These may include the user ID or a shopping cart ID. Code that is well-instrumented for Observability allows you to answer complex questions that are easy to miss when examining aggregate performance.

Observability and monitoring are essential practices in modern software development and operations. While observability focuses on understanding system behaviour through comprehensive data collection and analysis, monitoring uses predefined metrics to assess performance and generate alerts.

By leveraging both approaches, organizations can gain a holistic view of their systems, enabling proactive measures, faster troubleshooting, and optimal performance. Embracing observability and monitoring as complementary practices can pave the way for more reliable, scalable, and efficient systems in the digital era.

Interactive Observability Lab 02

Telemetry Correlation Investigation

A symptom rarely tells you the root cause. Correlate metrics, logs, traces and network context to move from “the service is slow” to a defensible engineering diagnosis.

Incident Scenario

Checkout API latency has increased

The checkout service is reporting elevated p95 latency. CPU and memory appear normal. Your task is to inspect the telemetry signals and determine what is actually causing the degradation.

Metric

Service Metrics

p95 latency: 1.84s
CPU: 42%
Memory: 58%
Confirms the symptom, but does not explain the cause.
Logs

Application Logs

DB query timeout
retry_count=2
pool_wait=640ms
Suggests requests are waiting on a downstream dependency.
Trace

Distributed Trace

checkout 1.84s
→ payment 110ms
→ DB 1.46s
Shows where the request spent most of its time.
Network

Network Context

DB RTT: 22ms
Packet loss: 0%
Connections: 98%
Network path is healthy; connection pressure is significant.
Correlated Evidence
13:42:01  checkout-api  p95 latency        1.84s
13:42:02  checkout-api  CPU                42%
13:42:03  checkout-api  memory             58%
13:42:04  checkout-api  DB timeout         640ms
13:42:05  trace         checkout span      1.84s
13:42:05  trace         DB span            1.46s
13:42:06  database      connection pool    98%
13:42:07  network       packet loss        0%
13:42:07  network       RTT                22ms
↓
CORRELATION

Symptom:
High checkout latency

Not supported:
CPU saturation
Memory exhaustion
Network packet loss

Strong evidence:
DB span = 1.46s
DB timeout/retries
Connection pool = 98%

Likely bottleneck:
Database connection pressure / pool exhaustion
✓ Diagnosis confirmed

The application hosts are not CPU- or memory-bound and the network path is healthy. The trace identifies the database span as the dominant contributor to latency, while logs and connection-pool telemetry provide supporting evidence. The investigation therefore points to database connection pressure rather than an application-host or network failure.

Evidence inspected: 0 / 4
Engineering takeaway Observability becomes powerful when telemetry is correlated rather than viewed in isolation. Metrics tell you that something changed; logs provide events and context; traces show where time was spent; network telemetry helps eliminate or confirm infrastructure causes. The goal is not more telemetry — it is better evidence.

Understanding Monitoring & Observability

A: Understanding Monitoring: Monitoring collects data and metrics from a system to track its health and performance. It involves setting up various tools and agents that continuously observe and report on predefined parameters. These parameters include resource utilization, response times, and error rates. Monitoring provides real-time insights into the system’s behavior and helps identify potential issues or bottlenecks.

B: Unveiling Observability: Observability goes beyond traditional monitoring by understanding the system’s internal state and cause-effect relationships. It aims to provide a holistic view of the system’s behavior, even in unexpected scenarios. Observability encompasses three main pillars: logs, metrics, and traces. Logs capture detailed events and activities, metrics quantify system behavior over time, and traces provide end-to-end transaction monitoring. By combining these pillars, observability enables deep system introspection and efficient troubleshooting.

C: The Power of Contextual Insights: One of the key advantages of observability is its ability to provide contextual insights. Traditional monitoring may alert you when a specific metric exceeds a threshold, but it often lacks the necessary context to debug complex issues. With its comprehensive data collection and correlation capabilities, Observability allows engineers to understand the context surrounding a problem. Contextual insights help in root cause analysis, reducing mean time to resolution and improving overall system reliability.

D: The Role of Automation: Automation plays a crucial role in monitoring and observability. In monitoring, automation can help set up alerts, generate reports, and scale the monitoring infrastructure. On the other hand, observability requires automated instrumentation and data collection to handle the vast amount of information generated by modern systems. Automation enables engineers to focus on analyzing insights rather than spending excessive time on data collection and processing.

**Observability: The First Steps**

The first step towards achieving modern observability is to gather metrics, traces, and logs. From the collected data points, observability aims to generate valuable outcomes for decision-making. The decision-making process goes beyond resolving problems as they arise. Next-generation observability goes beyond application remediation, focusing on creating business value to help companies achieve their operational goals. This decision-making process can be enhanced by incorporating user experience, topology, and security data.

**Observability Platform**

A full-stack observability platform monitors every monitored host in your environment. Depending on the technologies used, an average of 500 metrics are generated per computational node. AWS, Azure, Kubernetes, and VMware Tanzu are some platforms that use observability to collect important key performance metrics for services and real-user monitored applications. 

Within a microservices environment, dozens, if not hundreds, of microservices call one another. Distributed tracing can help you understand how the different services connect and how your requests flow. 

**Pillars of Observability**

The three pillars of observability form a strong foundation for making data-driven decisions, but there are opportunities to extend observability. User experience and security details must be considered to gain a deeper understanding. A holistic, context-driven approach to advanced observability enables proactively addressing potential problems before they arise.

**The Role of Monitoring**

To understand the difference between observability and monitoring, we need first to discuss the role of monitoring. Monitoring is the evaluation that helps identify the most practical and efficient use of resources. So, the big question I put to you is what to monitor. This is the first step to preparing a monitoring strategy.

To fully understand if monitoring is enough or if you need to move to an observability platform, ask yourself a couple of questions. Firstly, consider what you should be monitoring, why you should be monitoring it, and how you should be monitoring it. 

Observability vs Monitoring

Observability tells you why something is happening. Monitoring tells you what is happening.

This is one of the most important distinctions in modern systems — especially for cloud, microservices, Kubernetes, SD‑WAN, and SASE environments. Below is a structured, high‑clarity breakdown you can drop directly into your architecture notes or blog.

Concise Takeaway

Monitoring detects known problems using predefined dashboards and alerts. Observability lets you ask new questions and debug unknown problems using correlated metrics, logs, and traces.

Core Difference

Monitoring

  • Answers known questions

  • Uses predefined alerts

  • Focuses on symptoms

  • Works well for static systems

  • Example: “CPU > 80% — alert.”

Observability

  • Answers unknown questions

  • Uses rich telemetry (metrics, logs, traces)

  • Focuses on root cause

  • Works for dynamic, distributed systems

  • Example: “Why did latency spike only for users in Dublin?”

Key Components Compared

Monitoring Components

  • Dashboards

  • Threshold Alerts

  • Health Checks

  • Basic Metrics (CPU, memory, disk)

Monitoring is about detecting failures.

Observability Components

  • Metrics

  • Logs

  • Traces

  • Correlation

  • Context Propagation

  • SLOs

Observability is about understanding system behavior.

Why Observability Is More Powerful

Observability allows engineers to:

  • debug issues without adding new dashboards

  • trace requests across microservices

  • correlate identity, network, and application behavior

  • understand cloud and SaaS performance

  • detect anomalies before users notice

  • support auto‑scaling and SLO‑driven reliability

  • troubleshoot SD‑WAN and SASE path issues

Monitoring cannot do this because it only sees surface‑level symptoms.

Modern Architecture Context

In Microservices

Monitoring: “Service A is slow.” Observability: “Service A is slow because Service B’s database is overloaded.”

In Kubernetes

Monitoring: “Pod restart count is high.” Observability: “Pod restarts are caused by node pressure triggered by a spike in p99 latency.”

In SD‑WAN / SASE

Monitoring: “Packet loss detected.” Observability: “Packet loss is caused by ISP congestion between Dublin and London, impacting Zoom traffic.”

In Cloud

Monitoring: “API error rate increased.” Observability: “API errors correlate with IAM permission failures after a misconfigured deployment.”

Outcome

Observability delivers:

  • faster root‑cause analysis

  • better reliability

  • stronger SLO enforcement

  • improved user experience

  • deeper insight into distributed systems

  • intelligent auto‑scaling decisions

  • unified visibility across cloud, SD‑WAN, and SASE

Monitoring alone is no longer enough — observability is the foundation of modern operations.

Observability & Service Mesh

What is a Cloud Service Mesh?

A Cloud Service Mesh is a design pattern that helps manage and secure microservices interactions. Essentially, it acts as a dedicated layer for controlling the network traffic between microservices. By introducing a service mesh, developers can offload much of the responsibility for service communication from the application code itself, making the entire system more resilient and easier to manage.

### Key Benefits of Implementing a Cloud Service Mesh

#### Enhanced Security

One of the primary advantages of a Cloud Service Mesh is the enhanced security it offers. With features like mutual TLS (mTLS) for encrypting communications between services, a service mesh ensures that data is protected as it travels through the network. This is particularly important in a multi-cloud or hybrid cloud environment where services might be spread across different platforms.

#### Improved Observability

Observability is another critical benefit. A Cloud Service Mesh provides granular insights into service performance, helping developers identify and troubleshoot issues quickly. Metrics, logs, and traces are collected systematically, offering a comprehensive view of the entire microservices ecosystem.

#### Traffic Management

Managing traffic between services becomes significantly easier with a Cloud Service Mesh. Features like load balancing, traffic splitting, and failover mechanisms are built-in, ensuring that service-to-service communication remains efficient and reliable. This is particularly beneficial for applications requiring high availability and low latency.

### Popular Cloud Service Mesh Solutions

Several solutions have emerged as leaders in the Cloud Service Mesh space. Istio, Linkerd, and Consul are among the most popular options, each offering unique features and benefits. Istio, for example, is known for its robust policy enforcement and telemetry capabilities, while Linkerd is praised for its simplicity and performance. Consul, on the other hand, excels in multi-cloud environments, providing seamless service discovery and configuration.

### Challenges and Considerations

While the benefits are compelling, implementing a Cloud Service Mesh is not without its challenges. Complexity can be a significant hurdle, particularly for organizations new to microservices architecture. The additional layer of infrastructure requires careful planning and management. Moreover, there is a learning curve associated with configuring and maintaining a service mesh, which can impact development timelines.

Example Product: Cisco AppDynamics

### Real-Time Monitoring: Keeping an Eye on Your Applications

One of the standout features of Cisco AppDynamics is its real-time monitoring capabilities. By continuously tracking the performance of your applications, AppDynamics provides instant insights into any issues that may arise. This allows businesses to quickly identify and address performance bottlenecks, ensuring that their applications remain responsive and reliable. Whether it’s tracking transaction times, monitoring server health, or keeping an eye on user interactions, Cisco AppDynamics provides a comprehensive view of your application’s performance.

### Advanced Analytics: Turning Data into Actionable Insights

Data is the lifeblood of modern businesses, and Cisco AppDynamics excels at turning raw data into actionable insights. With its advanced analytics engine, AppDynamics can identify patterns, trends, and anomalies in your application’s performance data. This empowers businesses to make informed decisions, optimize their applications, and proactively address potential issues before they impact users. From root cause analysis to predictive analytics, Cisco AppDynamics provides the tools you need to stay ahead of the curve.

### Comprehensive Diagnostics: Troubleshooting Made Easy

When performance issues do arise, Cisco AppDynamics makes troubleshooting a breeze. Its comprehensive diagnostics capabilities allow you to drill down into every aspect of your application’s performance. Whether it’s identifying slow database queries, pinpointing code-level issues, or tracking down problematic user interactions, AppDynamics provides the detailed information you need to resolve issues quickly and efficiently. This not only minimizes downtime but also ensures a seamless user experience.

### Enhancing User Experiences: The Ultimate Goal

At the end of the day, the ultimate goal of any application is to provide a positive user experience. Cisco AppDynamics helps businesses achieve this by ensuring that their applications are always performing at their best. By providing real-time monitoring, advanced analytics, and comprehensive diagnostics, AppDynamics enables businesses to deliver fast, reliable, and engaging applications that keep users coming back for more. In a competitive digital landscape, this can be the difference between success and failure.

Google Cloud Monitoring

Example: What is Ops Agent?

Ops Agent is a lightweight, flexible monitoring agent explicitly designed for Compute Engine instances. It allows you to collect and analyze essential metrics and logs from your virtual machines, providing valuable insights into your infrastructure’s health, performance, and security.

To start monitoring your Compute Engine instance with Ops Agent, you must install and configure it properly. The installation process is straightforward and can be done through the Google Cloud Console or the command line. Once installed, you can configure Ops Agent to collect specific metrics and logs based on your requirements.

Ops Agent offers a wide range of metrics and logs that can be collected and monitored. These include system-level metrics like CPU and memory usage, network traffic, disk I/O, and more. Additionally, Ops Agent allows you to gather application-specific metrics and logs, providing deep insights into the performance and behavior of your applications running on the Compute Engine instance.

Options: Open source or commercial

Knowing this lets you move into the different tools and platforms available. Some of these tools will be open source, and others commercial. When evaluating these tools, one word of caution: does each tool work in a silo, or can it be used across technical domains? Silos are breaking agility in every form of technology.

 

Interactive Observability Lab 03

Four Golden Signals + Alerting Lab

Use the four Golden Signals — latency, traffic, errors and saturation — to decide whether a service needs attention and whether an alert is actionable.

Production Scenario

You are operating an API platform during a traffic increase. Review the current signals, choose the appropriate alert sensitivity, and decide whether the condition deserves an immediate engineering response.

L

Latency

1.72s
Time required to serve requests. Rising latency can expose customer-visible degradation.
Elevated
T

Traffic

8.4k r/s
Demand entering the system. Traffic spikes can explain changes in other signals.
Normal
E

Errors

3.8%
Failed requests indicate whether users are receiving unsuccessful outcomes.
Critical
S

Saturation

91%
How close a constrained resource is to its usable capacity.
High
Alert Evaluation
ACTION REQUIRED
Error rate is above the example critical threshold and saturation is high. This condition is potentially customer-impacting.

Decision score: 0 / 1
Engineering takeaway The Golden Signals provide a compact view of service health, but an alert should represent a condition that someone can act on. Alerting on every small metric fluctuation creates noise; alerting only on infrastructure thresholds can miss customer impact. Good alerting connects telemetry to meaningful service behaviour and operational action.
OpenShift Security Context Constraints

OpenShift Security Best Practices

OpenShift Security Best Practices

Security is a fundamental requirement for modern container platforms, particularly when OpenShift is used to host business-critical applications and services. Securing an OpenShift environment requires more than protecting individual containers. The security model spans identities, workloads, images, namespaces, network traffic, cluster configuration, secrets, observability, and the underlying infrastructure.

OpenShift provides a number of security controls through its integration with Kubernetes, including role-based access control, security contexts, network policies, image controls, secrets management, admission mechanisms, and platform-level security features. The challenge for engineering teams is to combine these capabilities into a coherent security architecture rather than relying on a single defensive layer.

Container and Image Security:

Container security begins before a workload reaches the cluster. Images should be sourced from trusted registries, scanned for known vulnerabilities, and controlled through an appropriate image lifecycle. Teams should minimise unnecessary packages, avoid embedding credentials in images, use immutable image references where appropriate, and rebuild images when important vulnerabilities are identified. Runtime controls should also enforce least privilege and prevent containers from receiving unnecessary Linux capabilities or elevated permissions.

Identity and Access Control:

Access to OpenShift should follow the principle of least privilege. Role-Based Access Control determines which users and service accounts can perform specific actions within the cluster. Permissions should be assigned through appropriate roles and groups rather than broad administrative access. Authentication should integrate with the organisation’s identity provider, while privileged accounts and service accounts should be reviewed regularly to reduce unnecessary access.

Workload Security:

OpenShift workloads should run with secure security contexts and only the privileges they require. Controls around user IDs, Linux capabilities, privilege escalation, filesystem access, and host interaction help reduce the impact of a compromised workload. OpenShift’s Security Context Constraints can be used to control the security characteristics under which pods are permitted to run. These controls are particularly important because containers share the underlying host kernel and therefore should not be treated as completely isolated virtual machines.

Network Security:

Network segmentation limits the ability of a compromised workload to communicate with unrelated services. Kubernetes NetworkPolicy can be used to define permitted traffic between namespaces, pods, and external destinations. In OpenShift environments, network architecture should also consider ingress and egress paths, service exposure, cluster infrastructure, and any additional network controls provided by the platform. A secure design should make allowed communication explicit rather than assuming that workloads require unrestricted connectivity.

Secrets and Sensitive Data:

Credentials, API keys, certificates, and other sensitive information should not be embedded directly into container images or application source code. OpenShift Secrets provide a mechanism for managing sensitive configuration, although organisations should also consider encryption, access controls, rotation procedures, external secret-management platforms, and the risks associated with exposing secrets to workloads. Secret access should be restricted to the applications and service accounts that genuinely require it.

Monitoring, Logging, and Detection:

Security controls are significantly more effective when cluster activity can be observed and investigated. Centralised application and platform logs, audit events, metrics, and security telemetry can help identify unusual authentication activity, privilege changes, unexpected workload behaviour, and other indicators of compromise. Alerting should focus on meaningful security and operational conditions rather than generating large volumes of low-value notifications.

Platform Updates and Vulnerability Management:

Keeping OpenShift and its underlying components current is an important part of maintaining the cluster’s security posture. Organisations should maintain a defined process for evaluating OpenShift releases, security advisories, container image vulnerabilities, operating-system updates, and dependencies. Updates should be tested and deployed through controlled processes while avoiding unnecessary exposure from unsupported or obsolete platform versions.

Security Policies and Admission Controls:

Security should be enforced as early as possible in the workload lifecycle. Policy and admission controls can prevent deployments that violate organisational requirements, such as running privileged workloads unnecessarily, using unapproved registries, or deploying images that fail defined security checks. This approach moves security from a purely reactive operational function towards preventative controls built into the application delivery process.

Securing OpenShift therefore requires a layered approach. Image security protects what enters the platform, identity and RBAC control who can operate it, workload security limits what applications can do, network policies restrict communication, secrets management protects sensitive information, and observability provides the evidence required to detect and investigate abnormal behaviour.

A strong OpenShift security posture is not created by a single configuration or security product. It is an ongoing engineering process that combines least privilege, secure workload design, policy enforcement, vulnerability management, monitoring, and regular platform maintenance. As applications and infrastructure evolve, these controls should be continuously reviewed to ensure that security remains aligned with the risks and operational requirements of the environment.

Highlights: OpenShift Security Best Practices

Interactive OpenShift Security Lab 01

OpenShift Security Posture Lab

Evaluate an OpenShift environment across multiple security layers. Decide whether each control represents a secure production posture or an avoidable security risk.

Scenario

Production OpenShift Cluster Review

A security team is reviewing a production cluster before approving a new workload. Examine the configuration below and identify which controls should be accepted and which require remediation.

Identity & Access
Secure
External identity provider is integrated. MFA is enforced and developers receive project-level roles rather than cluster-wide administrator privileges.
How would you classify this control?
Workload Privileges
Risk
A workload requests privileged execution, host networking and broad Linux capabilities even though the application only requires normal application-level access.
How would you classify this control?
Network Segmentation
Secure
NetworkPolicy rules restrict application communication to required services. Database workloads are not openly reachable by unrelated namespaces.
How would you classify this control?
Container Images
Risk
Production workloads pull images from an unapproved registry without vulnerability scanning or image-signing verification.
How would you classify this control?
Audit & Monitoring
Secure
API audit events are collected centrally and security-relevant activity is monitored for investigation and incident response.
How would you classify this control?
Secrets Management
Risk
Application credentials are embedded directly into container images and shared between multiple workloads instead of being managed separately.
How would you classify this control?
Security review progress 0 / 6 reviewed
Security Posture
0 / 6

Least privilege: avoid privileged workloads when application requirements do not justify them.
Supply chain: trust, scan and verify images before deployment.
Segmentation: restrict workload communication to required paths.
Identity: authenticate strongly and authorize with appropriate scope.
Engineering takeaway OpenShift security is not a single control. A strong posture combines identity and RBAC, workload restrictions such as SCCs, network segmentation, trusted images, secrets protection, and continuous auditing. The objective is to reduce both the probability of compromise and the blast radius when something goes wrong.

OpenShift Security 

**Understanding the Security Landscape**

Before diving into specific security measures, it’s essential to understand the overall security landscape of OpenShift. OpenShift is built on top of Kubernetes, inheriting its security features while also providing additional layers of protection. These include role-based access control (RBAC), network policies, and security context constraints (SCCs). Understanding how these components interact is crucial for building a secure environment.

**Implementing Role-Based Access Control**

One of the foundational elements of OpenShift security is Role-Based Access Control (RBAC). RBAC allows administrators to define what actions users and service accounts can perform within the cluster. By assigning roles and permissions carefully, you can ensure that each user only has access to the resources and operations necessary for their role, minimizing the risk of unauthorized actions or data breaches.

**Network Policies for Enhanced Security**

Network policies are another vital aspect of securing an OpenShift environment. These policies determine how pods within a cluster can communicate with each other and with external resources. By creating strict network policies, you can isolate sensitive workloads, control traffic flow, and prevent unauthorized access, thus reducing the attack surface of your applications.

**Security Context Constraints and Pod Security**

Security Context Constraints (SCCs) are used to define the permissions and access controls for pods. By setting up SCCs, you can restrict the capabilities of pods, such as preventing them from running as a root user. This minimizes the risk of privilege escalation attacks and helps maintain the integrity of your applications. Regularly reviewing and updating SCCs is a best practice for maintaining a secure OpenShift environment.

OpenShift Security Best Practices

1- Implementing Strong Authentication: One of the fundamental aspects of OpenShift security is ensuring robust authentication mechanisms are in place. Utilize features like OpenShift’s built-in OAuth server or integrate with external authentication providers such as LDAP or Active Directory. Enforce multi-factor authentication for added security and regularly review and update access controls to restrict unauthorized access.

2- Container Image Security: Container images play a vital role in OpenShift deployments, and their security should not be overlooked. Follow these best practices: regularly update base images to patch security vulnerabilities, scan images for known vulnerabilities using tools like Clair or Anchor, utilize trusted registries for image storage, and implement image signing and verification to ensure image integrity.

3- Network Segmentation and Policies: Proper network segmentation is crucial to isolate workloads and minimize the impact of potential security breaches. Leverage OpenShift’s network policies to define and enforce communication rules between pods and projects. Implement ingress and egress filtering to control traffic flow and restrict access to sensitive resources. Regularly review and update network policies to align with your evolving security requirements.

4- Logging and Monitoring: Comprehensive logging and monitoring are essential for effectively detecting and responding to security incidents. Enable centralized logging by leveraging OpenShift’s logging infrastructure or integrating with external log management solutions. Implement robust monitoring tools to track resource usage, detect abnormal behavior, and set up alerts for security-related events. Regularly review and analyze logs to identify potential security threats and take proactive measures.

OpenShift Security Best Practices

OpenShift security is built on a layered, Kubernetes‑native model that protects clusters, containers, workloads, CI/CD pipelines, and the underlying infrastructure. Because OpenShift is used for enterprise‑grade, multi‑tenant, production workloads, its security posture must combine Zero Trust, platform hardening, runtime protection, and continuous compliance.

Below is a structured, high‑clarity section you can drop directly into your blog or architecture notes.

1. Cluster & Platform Hardening

  • Control‑plane isolation — Keep masters separate from worker nodes.

  • Secure etcd — Encrypt etcd at rest and restrict access.

  • API server restrictions — Limit API exposure, enforce RBAC, and use audit logging.

  • Network segmentation — Separate workloads using namespaces and NetworkPolicies.

  • TLS everywhere — Enforce TLS for all internal and external communication.

OpenShift’s architecture already enforces many defaults, but tightening them is essential for production.

2. Identity & Access Security

  • RBAC — Use roles and role bindings instead of cluster‑wide permissions.

  • OAuth integration — Integrate with SSO (AD, LDAP, SAML, OIDC).

  • Least privilege — Avoid cluster‑admin for developers.

  • ServiceAccount hardening — Use dedicated service accounts per workload.

  • Secrets management — Store secrets in encrypted form and rotate regularly.

Identity is the new perimeter — OpenShift must enforce it strictly.

3. Container & Runtime Security

  • Image scanning — Scan images for CVEs before deployment.

  • Trusted registries — Block untrusted registries.

  • Pod Security Standards — Enforce restricted PSP‑equivalent policies.

  • SELinux enforcing — Mandatory access control for containers.

  • Runtime threat detection — Detect privilege escalation, crypto‑mining, and anomalous behavior.

OpenShift’s SELinux integration is one of its strongest security advantages.

4. Network Security

  • OpenShift SDN / OVN‑Kubernetes — Use NetworkPolicies to restrict pod‑to‑pod communication.

  • Ingress security — Enforce TLS, WAF, and rate limiting.

  • Egress control — Restrict outbound traffic to approved destinations.

  • Service Mesh — mTLS, identity‑aware routing, and policy enforcement.

Network segmentation is critical for multi‑tenant clusters.

5. Observability & Compliance

  • Audit logging — Track API calls, user actions, and cluster changes.

  • Metrics & traces — Use Prometheus, Grafana, and Jaeger.

  • Compliance Operator — Enforce CIS benchmarks and regulatory standards.

  • Cluster health monitoring — Detect drift, misconfigurations, and anomalies.

Observability is essential for detecting early signs of compromise.

6. CI/CD Pipeline Security

  • Secure build pipelines — Scan code, dependencies, and images.

  • Image signing — Use sigstore/cosign for trusted supply chain.

  • Admission controllers — Block deployments that violate security policies.

  • GitOps security — Protect ArgoCD/Flux credentials and enforce RBAC.

Supply chain security is now a mandatory requirement.

7. Node & OS Security

  • RHCOS hardening — Immutable OS reduces attack surface.

  • Node updates — Apply patches automatically using MachineConfig.

  • Kernel security — Use seccomp, SELinux, and cgroups.

  • Disable unnecessary services — Reduce exposure.

Nodes are the foundation — securing them protects the entire cluster.

Outcome

Implementing OpenShift security best practices delivers:

  • hardened cluster and control plane

  • secure workloads and containers

  • strong identity enforcement

  • Zero Trust network segmentation

  • protected CI/CD pipelines

  • continuous compliance

  • runtime threat detection

  • reliable, production‑grade security posture

OpenShift becomes a secure, resilient, enterprise‑ready Kubernetes platform.

Understanding Cluster Access

To begin our journey, let’s establish a solid understanding of cluster access. Cluster access refers to authenticating and authorizing users or entities to interact with an Openshift cluster. It involves managing user identities, permissions, and secure communication channels.

  • Implementing Multi-Factor Authentication (MFA)

Multi-factor authentication (MFA) requires users to provide multiple forms of identification, adding an extra layer of security. Enabling MFA within your Openshift cluster can significantly reduce the risk of unauthorized access. This section will outline the steps to configure and enforce MFA for enhanced cluster access security.

  • Role-Based Access Control (RBAC)

RBAC is a crucial component of Openshift security, allowing administrators to define and manage user permissions at a granular level. We will explore the concept of RBAC and its practical implementation within an Openshift cluster. Discover how to define roles, assign permissions, and effectively control access to various resources.

  • Secure Communication Channels

Establishing secure communication channels is vital to protect data transmitted between cluster components. In this section, we will discuss the utilization of Transport Layer Security (TLS) certificates to encrypt communication and prevent eavesdropping or tampering. Learn how to generate and manage TLS certificates within your Openshift environment.

  • Continuous Monitoring and Auditing

Maintaining a robust security posture involves constantly monitoring and auditing cluster access activities. Through integrating monitoring tools and auditing mechanisms, administrators can detect and respond to potential security breaches promptly. Uncover the best practices for implementing a comprehensive monitoring and auditing strategy within your Openshift cluster.

  • Threat modelling

A threat model maps out the likelihood and impact of potential threats to your system. Your security team is busy evaluating the risks to your platform, so it is essential to think about and evaluate them.

OpenShift clusters are no different, so when hardening your cluster, keep that in mind. Using the Kubeadmin user is probably fine if you use CodeReadContainers on your laptop to learn OpenShift. Access control and RBAC rules are probably a good idea for your company’s production clusters exposed to the internet.

If you model threats beforehand, you can explain what you did to protect your infrastructure and why you may not have taken specific other actions.

**Stricter Security than Kubernetes**

OpenShift has stricter security policies than Kubernetes. For instance, running a container as root is forbidden. To enhance security, it also offers a secure-by-default option. Kubernetes doesn’t have built-in authentication or authorization capabilities, so developers must manually create bearer tokens and other authentication procedures.

OpenShift provides a range of security features, including role-based access control (RBAC), image scanning, and container isolation, that help ensure the safety of containerized applications.

 

Interactive OpenShift Security Lab 02

SCC & Pod Security Investigation

Investigate a workload requesting excessive privileges. Identify which security controls increase the attack surface and decide whether the workload should be admitted.

Incident scenario

“The application needs elevated privileges.”

A development team has submitted a production deployment. They claim the application requires broad host access, but the security team suspects several requests are unnecessary. Review the workload configuration before allowing it to run.

Workload security context

securityContext:
  privileged: true
  allowPrivilegeEscalation: true
  runAsUser: 0
  readOnlyRootFilesystem: false
  capabilities:
    add: ["SYS_ADMIN"]

hostNetwork: true
hostPID: true
Privileged mode
Removes important isolation boundaries and should be exceptional.
High risk
Root UID
Running as root increases the impact of a successful container compromise.
Review
SYS_ADMIN
A broad Linux capability with significant security implications.
High risk
Host namespaces
Host network and PID access reduce workload isolation.
High risk

Investigation decision

Should this workload be admitted with the requested security context?
Investigation progress 0 / 3 findings
Which remediation is most appropriate?
Investigation result
0 / 3

Engineering principle: SCCs govern what security contexts workloads are permitted to use. They are not a reason to grant privileged access simply because an application requests it. Start with the application's actual requirements and reduce privileges wherever possible.
Engineering takeaway Container security is fundamentally about reducing privilege and preserving isolation. In OpenShift, evaluate privileged mode, UID strategy, Linux capabilities, host namespaces, privilege escalation and other workload security settings together. A secure design asks “what does this application actually need?” rather than “which policy makes it run?”

OpenShift Security Best Practices

Securing containerized environments

Securing containerized environments is considerably different from securing the traditional monolithic application because of the inherent nature of the microservices architecture. We went from one to many, and there is a clear difference in attack surface and entry points. So, there is much to consider for OpenShift network security and OpenShift security best practices, including many Docker security options.

The application stack previously had very few components, maybe just a cache, web server, and database separated and protected by a context firewall. The most common network service allows a source to reach an application, and the sole purpose of the network is to provide endpoint reachability.

As a result, the monolithic application has few entry points, such as ports 80 and 443. Not every monolithic component is exposed to external access and must accept requests directly, so we designed our networks around these facts. The following diagram provides information on the threats you must consider for container security.

container security
Diagram: Container security. Source Neuvector

Container Security Best Practices

1. Secure Authentication and Authorization: One of the fundamental aspects of OpenShift security is ensuring that only authorized users have access to the platform. Implementing robust authentication mechanisms, such as multifactor authentication (MFA) or integrating with existing identity management systems, is crucial to prevent unauthorized access. Additionally, defining fine-grained access controls and role-based access control (RBAC) policies will help enforce the principle of least privilege.

2. Container Image Security: OpenShift leverages containerization technology, which brings its security considerations. It is essential to use trusted container images from reputable sources and regularly update them to include the latest security patches. Implementing image scanning tools to detect vulnerabilities and malware within container images is also recommended. Furthermore, restricting privileged containers and enforcing resource limits will help mitigate potential security risks.

3. Network Security: OpenShift supports network isolation through software-defined networking (SDN). It is crucial to configure network policies to restrict communication between different components and namespaces, thus preventing lateral movement and unauthorized access. Implementing secure communication protocols, such as Transport Layer Security (TLS), between services and enforcing encryption for data in transit will further enhance network security.

4. Monitoring and Logging: A robust monitoring and logging strategy is essential for promptly detecting and responding to security incidents. OpenShift provides built-in monitoring capabilities, such as Prometheus and Grafana, which can be leveraged to monitor system health, resource usage, and potential security threats. Additionally, enabling centralized logging and auditing of OpenShift components will help identify and investigate security events.

5. Regular Vulnerability Assessments and Penetration Testing: To ensure the ongoing security of your OpenShift environment, it is crucial to conduct regular vulnerability assessments and penetration testing. These activities will help identify any weaknesses or vulnerabilities within the platform and its associated applications. Addressing these vulnerabilities promptly will minimize the risk of potential attacks and data breaches.

**OpenShift Security**

OpenShift delivers all the tools you need to run software on top of it with SRE paradigms, from a monitoring platform to an integrated CI/CD system that you can use to monitor and run both the software deployed to the OpenShift cluster and the cluster itself. So, the cluster and the workload that runs in it need to be secured. 

From a security standpoint, OpenShift provides robust encryption controls to protect sensitive data, including platform secrets and application configuration data. In addition, OpenShift optionally utilizes FIPS 140-2 Level 1 compliant encryption modules to meet security standards for U.S. federal departments.

This post highlights OpenShift security and provides security best practices and considerations when planning and operating your OpenShift cluster. These will give you a starting point. However, as clusters and bad actors are ever-evolving, it is important to revise the steps you took.

**Central security architecture**

Therefore, we often see security enforcement in a fixed central place in the network infrastructure. This could be, for example, a significant security stack consisting of several security appliances. We are often referred to as a kludge of devices. As a result, the individual components within the application need not worry about carrying out any security checks as they occur centrally for them.

On the other hand, with the common microservices architecture, those internal components are specifically designed to operate independently and accept requests alone, which brings considerable benefits to scaling and deploying pipelines.

However, each component may now have entry points and accept external connections. Therefore, they need to be concerned with security individually and not rely on a central security stack to do this for them.

**The different container attack vectors** 

These changes have considerable consequences for security and how you approach your OpenShift security best practices. The security principles still apply, and we still are concerned with reducing the blast radius, least privileges, etc. Still, they must be used from a different perspective and to multiple new components in a layered approach. Security is never done in isolation.

So, as the number of entry points to the system increases, the attack surface broadens, leading us to several docker container security attack vectors not seen with the monolithic. We have, for example, attacks on the Host, images, supply chain, and container runtime. There is also a considerable increase in the rate of change for these types of environments; an old joke says that a secure application is an application stack with no changes.

Open The Door To Bad Actors

So when you change, you can open the door to a bad actor. Today’s application varies considerably a few times daily for an agile stack. We have unit and security tests and other safety tests that can reduce mistakes, but no matter how much preparation you do, there is a chance of a breach whenever there is a change.

So, environmental changes affect security and some alarming technical challenges to how containers run as default, such as running as root by default and with a disturbing amount of capabilities and privileges. The following image displays attack vectors that are linked explicitly to containers.

container attack vectors
Diagram: Container attack vectors. Source Adriancitu

Challenges with Securing Containers

  • Containers running as root

As you know, containers run as root by default and share the Kernel of the Host OS. The container process is visible from the Host, which is a considerable security risk when a container compromise occurs. When a security vulnerability in the container runtime arose, and a container escape was performed, as the application ran as root, it could become root on the underlying Host.

Therefore, if a bad actor gets access to the Host and has the correct privileges, it can compromise all the hosts’ containers.

  • Risky Configuration

Containers often run with excessive privileges and capabilities—much more than they need to do their job efficiently. As a result, we need to consider what privileges the container has and whether it runs with any unnecessary capabilities it does not need.

Some of a container’s capabilities may be defaults that fall under risky configurations and should be avoided. You should keep an eye on the CAP_SYS_ADMIN flag, which grants access to an extensive range of privileged activities.

  • Excessive container isolation

The container has isolation boundaries by default with namespace and control groups ( when configured correctly). However, granting excessive container capabilities will weaken the isolation between the container, this Host, and other containers on the same Host. This is essentially removing or dissolving the container’s ring-fence capabilities.

Starting OpenShift Security Best Practices

Then, we have security with OpenShift, which overcomes many of the default security risks you have with running containers. And OpenShift does much of this out of the box. If you want further information on securing an OpenShift cluster, kindly check out my course for Pluralsight on OpenShift Security and OpenShift Network Security.

OpenShift Container Platform (formerly known as OpenShift Enterprise) or OCP is Red Hat’s offering for the on-premises private platform (PaaS). OpenShift is based on the Origin open-source project and is a Kubernetes distribution.

The foundation of the OpenShift Container Platform and OpenShift Network Security is based on Kubernetes and, therefore, shares some of the same networking technology and some enhancements. However, as you know, Kubernetes is a complex beast and can be utilized by itself when trying to secure clusters.  OpenShift does an excellent job of wrapping Kubernetes in a layer of security, such as using Security Context Constraints (SCCs) that give your cluster a good security base.

**Security Context Constraints**

By default, OpenShift prevents the cluster container from accessing protected functions. These functions—Linux features such as shared file systems, root access, and some core capabilities such as the KILL command—can affect other containers running in the same Linux kernel, so the cluster limits access.

Most cloud-native applications work fine with these limitations, but some (especially stateful workloads) need greater access. Applications that require these functions can still use them but need the cluster’s permission.

The application’s security context specifies the permissions that the application needs, while the cluster’s security context constraints specify the permissions that the cluster allows. An SC with an SCC enables an application to request access while limiting the access that the cluster will grant.

What are security contexts and security context constraints?

A pod configures a container’s access with permissions requested in the pod’s security context and approved by the cluster’s security context constraints:

-A security context (SC), defined in a pod, enables a deployer to specify a container’s permissions to access protected functions. When the pod creates the container, it configures it to allow these permissions and block all others. The cluster will only deploy the pod if the permissions it requests are permitted by a corresponding SCC.

-A security context constraint (SCC), defined in a cluster, enables an administrator to control pod permissions, which manage containers’ access to protected Linux functions. Similarly to how role-based access control (RBAC) manages users’ access to a cluster’s resources, an SCC manages pods’ access to Linux functions.

By default, a pod is assigned an SCC named restricted that blocks access to protected functions in OpenShift v4.10 or earlier. Instead, in OpenShift v4.11 and later, the restricted-v2 SCC is used by default. For an application to access protected functions, the cluster must make an SCC that allows it to be available to the pod.

SCC grants access to protection functions

While an SCC grants access to protected functions, each pod needing access must request it. To request access to the functions its application needs, a pod specifies those permissions in the security context field of the pod manifest. The manifest also specifies the service account that should be able to grant this access.

When the manifest is deployed, the cluster associates the pod with the service account associated with the SCC. For the cluster to deploy the pod, the SCC must grant the permissions that the pod requests.

One way to envision this relationship is to think of the SCC as a lock that protects Linux functions, while the manifest is the key. The pod is allowed to deploy only if the key fits.

Security Context Constraints
Diagram: Security Context Constraints. Source is IBM

A final note: Security context constraint

When your application is deployed to OpenShift in a virtual data center design, the default security model will enforce that it is run using an assigned Unix user ID unique to the project for which you are deploying it. Now, we can prevent images from being run as the Unix root user. When hosting an application using OpenShift, the user ID that a container runs as will be assigned based on which project it is running in.

Containers cannot run as the root user by default—a big win for security. SCC also allows you to set different restrictions and security configurations for PODs.

So, instead of allowing your image to run as the root, which is a considerable security risk, you should run as an arbitrary user by specifying an unprivileged USER, setting the appropriate permissions on files and directories, and configuring your application to listen on unprivileged ports.

OpenShift Network Security

SCC defaults access:

Security context constraints let you drop privileges by default, which is essential and still the best practice. Red Hat OpenShift security context constraints (SCCs) ensure that no privileged containers run on OpenShift worker nodes by default—another big win for security. Access to the host network and host process IDs are denied by default. Users with the required permissions can adjust the default SCC policies to be more permissive.

So, when considering SCC, consider SCC admission controllers as restricting POD access, similar to how RBAC restricts user access. To control the behavior of pods, we have security context constraints (SCCs). These cluster-level resources define what resources pods can access and provide additional control. 

Security context constraints let you drop privileges by default, which is a critical best practice. With Red Hat OpenShift SCCs, no privileged containers run on OpenShift worker nodes. Access to the host network and host process IDs is denied by default, a big win for OpenShift security.

Restricted security context constraints (SCCs):

A few SCCs are available by default, and you may have the head of the restricted SCC. By default, all pods, except those for builds and deployments, use a default service account assigned by the restricted SCC, which doesn’t allow privileged containers – that is, those running under the root user and listening on privileged ports are ports under <1024. SCC can be used to manage the following:

    1. Privilege Mode: This setting allows or denies a container from running in privilege mode. As you know, privilege mode bypasses any restriction such as control groups, Linux capabilities, secure computing profiles, 
    2. Privilege Escalation: This setting enables or disables privilege escalation inside the container ( all privilege escalation flags)
    3. Linux Capabilities: This setting allows the addition or removal of specific Linux capabilities
    4. Seccomp profile – this setting shows which secure computing profiles are used in a pod.
    5. Root-only file system: this makes the root file system read-only 

The goal is to assign the fewest possible capabilities for a pod to function fully. This least-privileged model ensures that pods can’t perform tasks on the system that aren’t related to their application’s proper function. The default value for the privileged option is False; setting the privileged option to True is the same as giving the pod the capabilities of the root user on the system. Although doing so shouldn’t be common practice, privileged pods can be helpful under certain circumstances. 

OpenShift Network Security: Authentication

Authentication refers to the process of validating one’s identity. Usually, users aren’t created in OpenShift but are provided by an external entity, such as the LDAP server or GitHub. The only part where OpenShift steps in is authorization—determining roles and permissions for a user.

OpenShift supports integration with various identity management solutions in corporate environments, such as FreeIPA/Identity Management, Active Directory, GitHub, Gitlab, OpenStack Keystone, and OpenID.

OpenShift Network Security: Users and identities

A user is any human actor who can request the OpenShift API to access resources and perform actions. Users are typically created in an external identity provider, usually a corporate identity management solution such as Lightweight Directory Access Protocol (LDAP) or Active Directory.

To support multiple identity providers, OpenShift relies on the concept of identities as a bridge between users and identity providers. A new user and identity are created upon the first login by default. There are four ways to map users to identities:

OpenShift Network Security: Service accounts

Service accounts allow us to control API access without sharing users’ credentials. Pods and other non-human actors use them to perform various actions and are a central vehicle by which their access to resources is managed. By default, three service accounts are created in each project:

OpenShift Network Security:  Authorization and role-based access control

Authorization in OpenShift is built around the following concepts:

Rules: Sets of actions allowed to be performed on specific resources.
Roles are collections of rules that allow them to be applied to a user according to a specific user profile. They can be used either at the cluster or project level.
Role bindings are associations between users or groups and roles. A given user or group can be associated with multiple roles.

If pre-defined roles aren’t sufficient, you can always create custom roles with just the specific rules you need.

Interactive OpenShift Security Lab 03

OpenShift Network Segmentation Lab

Build a simple application security boundary using NetworkPolicy. Decide which traffic should be permitted and then test whether the policy limits lateral movement.

Scenario

Protect the application tiers

Your namespace contains a frontend, an API service and a database. The API needs to reach the database, but the frontend should not connect directly to it. A compromised workload must also be prevented from freely moving across the namespace.

Application topology

F
Frontend
Public-facing application
port 443
→
A
API
Application service
port 8080
→
DB
Database
Internal service
port 5432
policy: api-to-database
target: database
allow: api → database:5432
deny: frontend → database
deny: unrelated → database
Lateral-movement test

Assume the frontend pod has been compromised. Can the attacker directly connect to the database service?

Traffic investigation

API → Database :5432
Expected: Allow
Required application traffic. The API needs database access.
Frontend → Database :5432
Expected: Deny
The frontend has no legitimate reason to connect directly to the database.
Unknown workload → Database :5432
Expected: Deny
Unrelated workloads should not automatically gain access to the database tier.
Frontend → API :8080
Expected: Allow
Required application-tier traffic.
Policy test progress 0 / 4 correct
Network security assessment
0 / 4

Engineering takeaway NetworkPolicy should express the application's required communication paths, not simply “open the namespace.” A useful security boundary allows necessary traffic while reducing unnecessary east-west connectivity. This limits lateral movement when a workload or credential is compromised.
System Observability

Distributed Systems Observability

Distributed Systems Observability

In the realm of modern technology, distributed systems have become the backbone of numerous applications and services. However, the increasing complexity of such systems poses significant challenges when it comes to monitoring and understanding their behavior. This is where observability steps in, offering a comprehensive solution to gain insights into the intricate workings of distributed systems. In this blog post, we will embark on a captivating journey into the realm of distributed systems observability, exploring its key concepts, tools, and benefits.

Observability, as a concept, enables us to gain deep insights into the internal state of a system based on its external outputs. When it comes to distributed systems, observability takes on a whole new level of complexity. It encompasses the ability to effectively monitor, debug, and analyze the behavior of interconnected components across a distributed architecture. By employing various techniques and tools, observability allows us to gain a holistic understanding of the system's performance, bottlenecks, and potential issues.

To achieve observability in distributed systems, it is crucial to focus on three interconnected components: logs, metrics, and traces.

Logs: Logs provide a chronological record of events and activities within the system, offering valuable insights into what has occurred. By analyzing logs, engineers can identify anomalies, track down errors, and troubleshoot issues effectively.

Metrics: Metrics, on the other hand, provide quantitative measurements of the system's performance and behavior. They offer a rich source of data that can be analyzed to gain a deeper understanding of resource utilization, response times, and overall system health.

Traces: Traces enable the visualization and analysis of transactions as they traverse through the distributed system. By capturing the flow of requests and their associated metadata, traces allow engineers to identify bottlenecks, latency issues, and performance optimizations.

In the ever-evolving landscape of distributed systems observability, a plethora of tools and frameworks have emerged to simplify the process. Prominent examples include:

1. Prometheus: A powerful open-source monitoring and alerting system that excels in collecting and storing metrics from distributed environments.

2. Jaeger: An end-to-end distributed tracing system that enables the visualization and analysis of transaction flows across complex systems.

3. ELK Stack: A comprehensive combination of Elasticsearch, Logstash, and Kibana, which collectively offer powerful log management, analysis, and visualization capabilities.

4. Grafana: A widely-used open-source platform for creating rich and interactive dashboards, allowing engineers to visualize metrics and logs in real-time.

The adoption of observability in distributed systems brings forth a multitude of benefits. It empowers engineers and DevOps teams to proactively detect and diagnose issues, leading to faster troubleshooting and reduced downtime. Observability also aids in capacity planning, resource optimization, and identifying performance bottlenecks. Moreover, it facilitates collaboration between teams by providing a shared understanding of the system's behavior and enabling effective communication.

In the ever-evolving landscape of distributed systems, observability plays a pivotal role in unraveling the complexity and gaining insights into system behavior. By leveraging the power of logs, metrics, and traces, along with robust tools and frameworks, engineers can navigate the intricate world of distributed systems with confidence. Embracing observability empowers organizations to build resilient, high-performing systems that can withstand the challenges of today's digital landscape.

Highlights: Distributed Systems Observability

Interactive Distributed Systems Lab 01

Distributed Failure Investigation Lab

A request is slow, but the frontend and API appear healthy. Investigate the dependency chain and determine where the partial failure actually originates.

SCENARIO

Checkout requests have become noticeably slower. Users report timeouts, while the frontend and API dashboards still look relatively normal. Your task is to correlate the service path and identify the most likely root cause.

Request dependency path — tap a service to inspect it

Evidence

Select telemetry evidence before choosing your diagnosis.

Investigation Console

Inspect the evidence, then select the most likely root cause.
Awaiting inspection
Signal Select evidence
Observation —
Implication —
What is the most likely root cause?
Evidence inspected: 0 / 4
Engineering takeaway: In a distributed system, the component showing the user-visible symptom is not necessarily the component causing the failure. Correlate latency, dependency health and network evidence across service boundaries to distinguish a local problem from a downstream partial failure.

The Role of Distributed Systems

A – Several decades ago, only a handful of mission-critical services worldwide were required to meet the availability and reliability requirements of today’s always-on applications and APIs. In response to user demand, every application must be built to scale nearly instantly to accommodate the potential for rapid, viral growth. Almost every app built today—whether a mobile app for consumers or a backend payment system—must meet these constraints and requirements.

B – Inherently, distributed systems are more reliable due to their distributed nature. When appropriately designed software engineers build these systems, they can benefit from more scalable organizational models. There is, however, a price to pay for these advantages.

C – Designing, building, and debugging these distributed systems can be challenging. A reliable distributed system requires significantly more engineering skills than a single-machine application, such as a mobile app or a web frontend. Regardless, distributed systems are becoming increasingly important. There is a corresponding need for tools, patterns, and practices to build them.

D – As digital transformation accelerates, organizations adopt multicloud environments to drive secure innovation and achieve speed, scale, and agility. As a result, technology stacks are becoming increasingly complex and scalable. Today, even the most straightforward digital transaction is supported by an array of cloud-native services and platforms delivered by various providers. To improve user experience and resilience, IT and security teams must monitor and manage their applications.

**Key Components of Observability**

Observability in distributed systems typically relies on three pillars: logs, metrics, and traces. Logs provide detailed records of events within the system, offering context for debugging issues. Metrics offer quantitative data, such as CPU usage and request rates, allowing teams to monitor system health and performance over time. Traces enable the tracking of requests as they move through the system, helping to pinpoint where latency or failures occur. Together, these components create a comprehensive picture of the system’s state and behavior.

**Challenges in Achieving Observability**

While observability is essential, achieving it in distributed systems is not without its challenges. The sheer volume of data generated by these systems can be overwhelming. Additionally, correlating data from disparate sources to form a cohesive narrative requires sophisticated tools and techniques. Moreover, ensuring that observability doesn’t introduce too much overhead or affect system performance is a delicate balancing act. Organizations must invest in the right infrastructure and expertise to tackle these challenges effectively.

**Best Practices for Enhancing Observability**

To maximize observability in distributed systems, organizations should adopt several best practices. Firstly, they should implement centralized logging and monitoring solutions that can aggregate data from all system components. Secondly, leveraging open standards like OpenTelemetry can facilitate consistent data collection and integration with various tools. Thirdly, incorporating automated alerting and anomaly detection can help teams proactively address issues before they impact users. Lastly, fostering a culture of collaboration between development and operations teams can ensure that observability is an ongoing, shared responsibility.

Distributed Systems Observability

Distributed systems observability is the discipline of understanding exactly what is happening inside large, multi‑component, multi‑region, microservice‑based architectures — without adding new code or dashboards every time something breaks. It is the evolution of monitoring, designed for systems where failures propagate across services, queues, caches, databases, networks, and cloud edges.

Below is a structured, high‑clarity section you can drop directly into your SRE, cloud, or microservices content.

Core Goal

Distributed systems observability allows engineers to reconstruct system behavior across dozens or hundreds of components using correlated telemetry:

  • Metrics

  • Logs

  • Traces

  • Events

  • Topology

It answers why something happened, not just what happened.

Key Challenges in Distributed Systems

Distributed systems introduce complexity that monitoring alone cannot handle:

  • partial failures

  • network partitions

  • retry storms

  • queue backpressure

  • cascading latency

  • multi‑region inconsistencies

  • service‑to‑service dependency chains

  • ephemeral workloads (Kubernetes)

Observability provides the visibility needed to debug these behaviors.

Core Components of Distributed Systems Observability

1. Metrics — High‑Level Health

Metrics show system symptoms across services:

  • latency (p95/p99)

  • error rate

  • throughput

  • queue depth

  • resource saturation

They reveal where a problem is occurring.

2. Logs — Detailed Forensics

Logs show:

  • state changes

  • errors

  • retries

  • timeouts

  • authentication failures

They reveal what happened inside each component.

3. Traces — End‑to‑End Request Flow

Traces show:

  • how a request moves across microservices

  • where bottlenecks occur

  • which dependency caused the slowdown

  • how long each span took

Traces reveal why the system behaved the way it did.

4. Topology Awareness

Distributed systems need dynamic topology visibility:

  • service maps

  • dependency graphs

  • network paths

  • cloud region routing

This reveals how everything is connected.

Distributed Systems Observability in Modern Architectures

Microservices

Observability correlates:

  • service latency

  • database bottlenecks

  • cache misses

  • downstream failures

  • retry loops

Kubernetes

Observability tracks:

  • pod lifecycle

  • node pressure

  • HPA/VPA scaling

  • container restarts

  • network overlay behavior

Cloud & Serverless

Observability reveals:

  • cold starts

  • throttling

  • multi‑region latency

  • API gateway behavior

  • IAM permission failures

Event‑Driven Systems

Observability shows:

  • queue depth

  • consumer lag

  • backpressure

  • dead‑letter queues

  • event replay patterns

Why Distributed Systems Need Observability

Distributed systems fail in non‑linear ways. Observability enables:

  • root‑cause analysis across multiple services

  • correlation of identity + network + application behavior

  • detection of cascading failures

  • SLO‑driven reliability

  • intelligent auto‑scaling

  • faster incident response

  • reduced MTTR

  • prevention of blind spots

Monitoring alone cannot do this.

📡 Tools for Distributed Systems Observability

  • OpenTelemetry — unified metrics, logs, traces.

  • Prometheus — metrics collection.

  • Grafana — dashboards and alerting.

  • Jaeger — distributed tracing.

  • Zipkin — tracing for microservices.

  • Splunk — analytics + correlation.

  • ThousandEyes — network + cloud path visibility.

  • Elastic — logs + metrics + traces.

Together, these form a complete observability pipeline.

Outcome

Distributed systems observability delivers:

  • complete insight across microservices

  • faster debugging of unknown failures

  • reliable auto‑scaling

  • strong SLO enforcement

  • reduced downtime

  • improved user experience

  • unified visibility across cloud, network, and application layers

It transforms distributed systems from opaque black boxes into fully observable, measurable, and debuggable architectures.

Cloud Service Mesh

### What is a Cloud Service Mesh?

A Cloud Service Mesh is a dedicated infrastructure layer that facilitates service-to-service communication in a microservices architecture. It abstracts the complex communication patterns between services into a manageable, secure, and observable framework. By deploying a service mesh, organizations can effectively manage the interactions of their microservices, ensuring seamless connectivity, security, and resilience.

### Key Benefits of Implementing a Cloud Service Mesh

1. **Enhanced Security**: A service mesh provides robust security features such as mutual TLS authentication, which encrypts communications between services. This ensures that data remains secure and tamper-proof as it travels across the network.

2. **Traffic Management**: With a service mesh, you can implement sophisticated traffic management policies, including load balancing, circuit breaking, and retries. This leads to improved performance and reliability of your distributed systems.

3. **Observability**: One of the standout features of a service mesh is its ability to provide deep observability into the interactions between services. Metrics, logs, and traces are collected and analyzed, offering invaluable insights into system health and performance.

### Enhancing Observability in Distributed Systems

Observability is a key concern in managing distributed systems. With the proliferation of microservices, tracking and understanding service interactions can become overwhelmingly complex. A Cloud Service Mesh addresses this challenge by offering comprehensive observability features:

– **Metrics Collection**: Collects real-time metrics on service performance, latency, error rates, and more.

– **Distributed Tracing**: Enables tracing of requests as they propagate through multiple services, helping identify bottlenecks and performance issues.

– **Centralized Logging**: Aggregates logs from various services, providing a unified view for easier troubleshooting and analysis.

These capabilities empower teams to detect issues early, optimize performance, and ensure the reliability of their applications.

### Real-World Applications and Use Cases

Several organizations have successfully implemented Cloud Service Meshes to transform their operations. For instance, financial institutions use service meshes to secure sensitive transactions, while e-commerce platforms leverage them to manage high traffic volumes during peak shopping seasons. By providing a robust framework for service communication, a service mesh enhances scalability, reliability, and security across industries.

Googles Ops Agent

Ops Agent is a lightweight agent that runs on your Compute Engine instances, collecting and forwarding metrics and logs to Google Cloud Monitoring and Logging. By installing Ops Agent on your instances, you gain real-time visibility into your Compute Engine’s performance and behavior.

To start monitoring your Compute Engine, you must install Ops Agent on your instances. The installation process is straightforward and can be done manually or through automation tools like Cloud Deployment Manager or Terraform. Once installed, the Ops Agent will automatically begin collecting metrics and logs from your Compute Engine.

Ops Agent allows you to customize the metrics and logs you want to monitor for your Compute Engine. Various options are available, allowing you to choose specific metrics and logs relevant to your application or system. By configuring metrics and logs, you can gain deeper insights and track the performance of critical components.

**Challenge: Fragmented Monitoring Tools**

Fragmented monitoring tools and manual analytics strategies challenge IT and security teams. The lack of a single source of truth and real-time insight makes it increasingly difficult for these teams to access the answers they need to accelerate innovation and optimize digital services. To gain insight, they must manually query data from various monitoring tools and piece together different sources of information.

This complex and time-consuming process distracts Team members from driving innovation and creating new value for the business and customers. In addition, many teams monitor only their mission-critical applications due to the effort involved in managing all these tools, platforms, and dashboards. The result is a multitude of blind spots across the technology stack, which makes it harder for teams to gain insights.

**Challenge: Kubernetes is Complex**

Understanding how Kubernetes adds to the complexity of technology stacks is imperative. In the drive toward modern technology stacks, it is the platform of choice for organizations refactoring their applications for the cloud-native world. Through dynamic resource provisioning, Kubernetes architectures can quickly scale services to new users and increase efficiency.

However, the constant changes in cloud environments make it difficult for IT and security teams to maintain visibility into them. To provide observability in their Kubernetes environments, these teams cannot manually configure various traditional monitoring tools. The result is that they are often unable to gain real-time insights to improve user experience, optimize costs, and strengthen security. Due to this visibility challenge, many organizations are delaying moving more mission-critical services to Kubernetes.

GKE-Native Monitoring

The Basics of GKE-Native Monitoring

GKE-Native Monitoring is a comprehensive monitoring solution provided by Google Cloud Platform (GCP) designed explicitly for GKE clusters. It offers deep insights into your applications’ performance and behavior, allowing you to proactively detect and resolve issues. With GKE-Native Monitoring, you can easily collect and analyze metrics, monitor logs, and set up alerts to ensure the reliability and availability of your applications.

One of the critical features of GKE-Native Monitoring is its ability to collect and analyze metrics from your GKE clusters. It provides preconfigured dashboards that display essential metrics such as CPU usage, memory utilization, and network traffic. Additionally, you can create custom dashboards tailored to your specific requirements, allowing you better to understand your application’s performance and resource consumption.

The Role of Megatrends

We have had a considerable drive with innovation that has spawned several megatrends that have affected how we manage and view our network infrastructure and the need for distributed systems observability. We have seen the decomposition of everything from one to many.

Many services and dependencies in multiple locations, aka microservices observability, must be managed and operated instead of the monolithic where everything is generally housed internally. The megatrends have resulted in a dynamic infrastructure with new failure modes not seen in the monolithic, forcing us to look at different systems observability tools and network visibility practices. 

Shift in Control

There has also been a shift in the point of control. As we move towards new technologies, many of these loosely coupled services or infrastructures your services depend on are not under your control. The edge of control has been pushed, creating different network and security perimeters. These parameters are now closer to the workload than a central security stack. Therefore, the workloads themselves are concerned with security.

Example Product: Cisco AppDynamics

### What is Cisco AppDynamics?

Cisco AppDynamics is an application performance management (APM) solution designed to provide real-time visibility into the performance of your applications. It helps IT professionals identify bottlenecks, diagnose issues, and optimize performance, ensuring a seamless user experience. With its powerful analytics, you can gain deep insights into your application stack, from the user interface to the backend infrastructure.

### Key Features and Capabilities

#### Real-Time Monitoring

One of the standout features of Cisco AppDynamics is its ability to monitor applications in real-time. This allows IT teams to detect and resolve issues as they occur, minimizing downtime and ensuring a smooth user experience. Real-time monitoring covers everything from user interactions to server performance, providing a comprehensive view of your application’s health.

#### End-User Experience Monitoring

Understanding how users interact with your application is crucial for delivering a high-quality experience. Cisco AppDynamics offers end-user experience monitoring, which tracks user sessions and interactions. This data helps you identify any pain points or performance issues that may be affecting user satisfaction.

#### Business Transaction Monitoring

Cisco AppDynamics takes a unique approach to monitoring by focusing on business transactions. By tracking the performance of individual transactions, you can gain a clearer understanding of how different parts of your application are performing. This level of granularity allows for more targeted optimizations and quicker issue resolution.

### Benefits of Using Cisco AppDynamics

#### Improved Application Performance

With its comprehensive monitoring and diagnostic capabilities, Cisco AppDynamics helps you identify and resolve performance issues quickly. This leads to faster load times, fewer errors, and an overall improved user experience.

#### Enhanced Operational Efficiency

By automating many of the monitoring and diagnostic processes, Cisco AppDynamics reduces the workload on your IT team. This allows them to focus on more strategic initiatives, driving greater value for your business.

#### Better Decision Making

The insights provided by Cisco AppDynamics enable better decision-making at all levels of your organization. Whether you’re looking to optimize resource allocation or plan for future growth, the data and analytics provided can inform your strategies and drive better outcomes.

### Integrations and Flexibility

Cisco AppDynamics offers seamless integrations with a wide range of third-party tools and platforms. Whether you’re using cloud services like AWS and Azure or CI/CD tools like Jenkins and GitHub, AppDynamics can integrate into your existing workflows, providing a unified view of your application’s performance.

 

Interactive Distributed Systems Lab 02

Distributed Trace Explorer

Follow one request across multiple services. Inspect individual spans, identify retries and downstream latency, and determine which component contributes most to the user-visible delay.

TRACE SCENARIO

Request checkout-7f31 takes approximately 1.82 seconds. The trace contains several downstream operations. Tap individual spans to inspect what happened inside the request.

Distributed trace — tap a span to inspect
1.82 s
1.49 s
430 ms
165 ms
1.12 s
Total trace duration
1.82 s
Database span
1.12 s
Retry detected
Yes

Span Inspector

Select a span to understand its role in the request.
No span selected
Service —
Duration —
Dependency —
Finding Tap a span

Trace Analysis

Choose the observation that best explains the trace.
Trace actions: 0 / 3
Engineering takeaway: A distributed trace turns one user request into a sequence of correlated spans. The key is not simply finding the longest span: engineers must understand dependency relationships, retries and whether work is sequential or overlapping before assigning causality.

Distributed Systems Observability

Distributed Systems

Today’s world of always-on applications and APIs has availability and reliability requirements that would have been needed of solely a handful of mission-critical services around the globe only a few decades ago. Likewise, the potential for rapid, viral service growth means that every application has to be built to scale nearly instantly in response to user demand.

Finally, these constraints and requirements mean that almost every application made—whether a consumer mobile app or a backend payments application—needs to be a distributed system. A distributed system is an environment where different components are spread across multiple computers on a network. These devices split up the work, harmonizing their efforts to complete the job more efficiently than if a single device had been responsible.

**The Key Components of Observability**

Observability in distributed systems is achieved through three main components: monitoring, logging, and tracing.

1. Monitoring: Monitoring involves continuously collecting and analyzing system metrics and performance indicators. It provides real-time visibility into the health and performance of the distributed system. By monitoring various metrics such as CPU usage, memory consumption, network traffic, and response times, engineers can proactively identify anomalies and make informed decisions to optimize system performance.

2. Logging: Logging involves recording events, activities, and errors within the distributed system. Log data provides a historical record that can be analyzed to understand system behavior and debug issues. Distributed systems generate vast amounts of log data, and effective log management practices, such as centralized log storage and log aggregation, are crucial for efficient troubleshooting.

3. Tracing: Tracing involves capturing the flow of requests and interactions between different distributed system components. It allows engineers to trace the journey of a specific request and identify potential bottlenecks or performance issues. Tracing is particularly useful in complex distributed architectures where multiple services interact.

**Benefits of Observability in Distributed Systems**

Adopting observability practices in distributed systems offers several benefits:

1. Enhanced Troubleshooting: Observability enables engineers to quickly identify and resolve issues by providing detailed insights into system behavior. With real-time monitoring, log analysis, and tracing capabilities, engineers can pinpoint the root cause of problems and take appropriate actions, minimizing downtime and improving system reliability.

2. Performance Optimization: By closely monitoring system metrics, engineers can identify performance bottlenecks and optimize system resources. Observability allows for proactive capacity planning and efficient resource allocation, ensuring optimal performance even under high loads.

3. Efficient Change Management: Observability facilitates monitoring system changes and their impact on overall performance. Engineers can track changes in metrics and easily identify any deviations or anomalies caused by updates or configuration changes. This helps maintain system stability and avoid unexpected issues.

What are VPC Flow Logs?

VPC Flow Logs is a feature offered by Google Cloud that captures and records network traffic information within Virtual Private Cloud (VPC) networks. Each network flow is logged, providing a comprehensive view of the traffic traversing. These logs include valuable information such as source and destination IP addresses, ports, protocol, and packet counts.

Once the VPC Flow Logs are enabled and data is being recorded, we can start leveraging the power of analysis. Google Cloud provides several tools and services for analyzing VPC Flow Logs. One such tool is BigQuery, a scalable and flexible data warehouse. By exporting VPC Flow Logs to BigQuery, we can perform complex queries, visualize traffic patterns, and detect anomalies using industry-standard SQL queries.

**How This Affects Failures**

The primary issue I have seen with my clients is that application failures are no longer predictable, and dynamic systems can fail creatively, challenging existing monitoring solutions. But, more importantly, the practices that support them. We have a lot of partial failures that are not just unexpected but not known or have never been seen before. For example, if you recall, we have the network hero. 

**The network Hero**

It is someone who knows every part of the network and has seen every failure at least once. These people are no longer helpful in today’s world and need proper Observation. When I was working as an Engineer, we would have plenty of failures, but more than likely, we would have seen them before. And there was a system in place to fix the error. Today’s environment is much different.

We can no longer rely on simply seeing a UP or Down, setting static thresholds, and then alerting based on those thresholds. A key point to note at this stage is that none of these thresholds considers the customer’s perspective.  If your POD runs at 80% CPU, does that mean the customer is unhappy?

When monitoring, you should look from your customer’s perspectives and what matters to them. Content Delivery Network (CDN) was one of the first to realize this game and measure what matters most to the customer.

Distributed Systems Observability

The different demands

So, the new, modern, and complex distributed systems place very different demands on your infrastructure and the people who manage it. For example, in microservices, there can be several problems with a particular microservice:

    • The microservices could be running under high resource utilization and, therefore, slow to respond, causing a timeout
    • The microservices could have crashed or been stopped and is, therefore, unavailable
    • The microservices could be fine, but there could be slow-running database queries.
    • So we have a lot of partial failures. 

Consequently, We can no longer predict

The significant shift we see with software platforms is that they evolve much quicker than the products and paradigms we use to monitor them. As a result, we need to consider new practices and technologies with dedicated platform teams and sound system observability. We can’t predict anything anymore, which puts the brakes on some traditional monitoring approaches, especially the metrics-based approach to monitoring.

I’m not saying that these monitoring tools are not doing what you want them to do. But, they work in a siloed environment, and there is a lack of connectivity. So we have monitoring tools working in silos in different parts of the organization and more than likely managed by other people trying to monitor a very dispersed application with multiple components and services in various places. 

Relying On Known Failures

Metric-Based Approach

A metrics-based monitoring approach relies on having previously encountered known failure modes. The metric-based approach relies on known failures and predictable failure modes. So, we have predictable thresholds that someone is considered to experience abnormal.

Monitoring can detect when these systems are either over or under the predictable thresholds that were previously set. Then, we can set alerts, and we hope that these alerts are actionable. This is only useful for variants of predictable failure modes.

Traditional metrics and monitoring tools can tell you any performance spikes or notice that a problem occurs. But they don’t let you dig into the source of the issues and let us slice and dice or see correlations between errors. If the system is complex, this approach is more challenging in getting to the root cause in a reasonable timeframe.

Google Cloud Trace

Example: Application Latency & Cloud Trace

Before we discuss Cloud Trace’s specifics, let’s establish a clear understanding of application latency. Latency refers to the time delay between a user’s action or request and the corresponding response from the application. It includes network latency, server processing time, and database query execution time. By comprehending the different factors contributing to latency, developers can proactively optimize their applications for improved performance.

Google Cloud Trace is a powerful diagnostic tool offered by Google Cloud Platform (GCP) that enables developers to identify and analyze application latency bottlenecks. It provides detailed insights into the flow of requests and events within an application, allowing developers to pinpoint areas of concern and optimize accordingly. Cloud Trace integrates seamlessly with other GCP services and provides a comprehensive view of latency across various components of an application stack.

Traditional style metrics systems

With traditional metrics systems, you had to define custom metrics, which were always defined upfront. This approach prevents us from starting to ask new questions about problems. So, it would be best to determine the questions to ask upfront.

Then, we set performance thresholds, pronounce them “good” or “bad, ” and check and re-check those thresholds. We would tweak the thresholds over time, but that was about it. This monitoring style has been the de facto approach, but we don’t now want to predict how a system can fail. Always observe instead of waiting for problems, such as reaching a certain threshold before acting.

**Metrics: Lack of connective event**

The metrics did not retain the connective event, so you cannot ask new questions in the existing dataset. These traditional system metrics could miss unexpected failure modes in complex distributed systems. Also, the condition detected via system metrics might be unrelated to what is happening.

An example of this could be an odd number of running threads on one component, which might indicate garbage collection is in progress or that slow response times are imminent in an upstream service.

**Users experience static thresholds**

User experience means different things to different sets of users. We now have a model where different service users may be routed through the system in other ways, using various components and providing experiences that can vary widely. We also know now that the services no longer tend to break in the same few predictable ways over and over.  

We should have a few alerts triggered by only focusing on symptoms that directly impact user experience and not because a threshold was reached.

The Challenge: Can’t reliably indicate any issues with user experience

If you use static thresholds, they can’t reliably indicate any issues with user experience. Alerts should be set up to detect failures that impact user experience. Traditional monitoring falls short in this regard. With traditional metrics-based monitoring, we rely on static thresholds to define optimal system conditions, which have nothing to do with user experience.

However, modern systems change shape dynamically under different workloads. Static monitoring thresholds can’t reflect impacts on user experience. They lack context and are too coarse.

Required: Distributed Systems Observability

Systems observability and reliability in distributed systems are practices. Rather than just focusing on a tool that logs, metrics, or alters, Observability is all about how you approach problems, and for this, you need to look at your culture. So you could say that Observability is a cultural practice that allows you to be proactive about findings instead of relying on a reactive approach that we are used to in the past.

Nowadays, we need a different viewpoint and want to see everything from one place. You want to know how the application works and how it interacts with the other infrastructure components, such as the underlying servers, physical or server, the network, and how data transfer looks in a transfer and stale state. 

Levels of Abstraction

What level of observation is needed to ensure everything performs as it should? What should you look at to obtain this level of detail?

Monitoring is knowing the data points and the entities from which we gather information. On the other hand, Observability is like putting all the data together. So monitoring is collecting data, and Observability is putting it together in one single pane of glass. Observability is observing the different patterns and deviations from the baseline; monitoring is getting the data and putting it into the systems. A vital part of an Observability toolkit is service level objectives (slos).

**Preference: Distributed Tracing**

We have three pillars of Systems Observability. There are Metrics, Traces, and Logging. So, defining or viewing Observability as having these pillars is an oversimplification. But for Observability, you need these in place. Observability is all about connecting the dots from each of these pillars.

If someone asked me which one I prefer, it would be distributed tracing. Distributed tracing allows you to visualize each step in service request executions. As a result, it doesn’t matter if services have complex dependencies. You could say that the complexity of the Dynamic systems is abstracted with distributed tracing.

**Use Case: Challenges without tracing**

For example, latency can stack up if a downstream database service experiences performance bottlenecks, resulting in high end-to-end latency. When latency is detected three or four layers upstream, it can be complicated to identify which component of the system is the root of the problem because now that same latency is being seen in dozens of other services.

**Distributed tracing: A winning formula**

Modern distributed systems tend to scale into a tangled knot of dependencies. Therefore, distributed tracing shows the relationships between various services and components in a distributed system. Traces help you understand system interdependencies. Unfortunately, those inter-dependencies can obscure problems and make them challenging to debug unless their relationships are clearly understood.

In distributed systems, observability is vital in ensuring complex architectures’ stability, performance, and reliability. Monitoring, logging, and tracing provide engineers with the tools to understand system behavior, troubleshoot issues, and optimize performance. By adopting observability practices, organizations can effectively manage their distributed systems and provide seamless and reliable services to their users.

Interactive Distributed Systems Lab 03

Observability Architecture & Dependency Lab

Distributed applications cross application, platform, network and external dependency boundaries. Explore the architecture and identify where an observability design can still leave engineers blind.

ARCHITECTURE SCENARIO

Your application runs on Kubernetes and uses several microservices, a service mesh, a database and an external payment provider. The team has application metrics and logs, but an incident still cannot be explained quickly. Find the visibility gaps.

Distributed application — tap components to inspect observability coverage
Application
User-facing workloads and service dependencies
→
Platform
Kubernetes and service-to-service communication
→
Dependencies
Data, network and third-party systems

Visibility Inspector

Inspect what telemetry exists and what may still be missing.
Select an architecture component
Component —
Telemetry —
Visibility —
Risk —

Investigation Tasks

Identify the three capabilities most important for explaining cross-service failures.
Architecture checks: 0 / 3
Engineering takeaway: Observability coverage should follow the dependency graph, not just the application code. A useful design connects telemetry across workloads, platform infrastructure, service-to-service communication, networks, databases and external dependencies so that failures can be correlated across boundaries.
Reliability in Distributed Systems

Reliability In Distributed System

Reliability In Distributed System

Distributed systems are now fundamental to modern computing. Cloud platforms, payment systems, e-commerce applications, content delivery networks, databases, telecommunications services, and AI platforms all depend on multiple machines and services working together across networks. This architecture provides scalability, flexibility, and resilience, but it also introduces a fundamental engineering problem: components can fail independently while the overall system is expected to continue operating.

Reliability in a distributed system is therefore not simply a question of whether individual servers are available. It is the ability of the overall system to continue providing correct and useful behaviour despite failures, delays, network problems, overloaded components, software defects, and unexpected operating conditions. A reliable architecture assumes that failures will occur and is designed to contain their impact rather than assuming that every component will always behave correctly.

Failure Models:

Distributed systems must account for different types of failure. A node may crash completely, a network connection may be interrupted, packets may be delayed or dropped, a dependency may become unavailable, or a service may respond much more slowly than expected. More complex failure models can also include partial failures, where one component remains operational while its communication with another component has failed. Designing around these failure modes is one of the central challenges of distributed-systems engineering.

Redundancy and Replication:

Redundancy reduces dependence on individual components. Data and services can be replicated across multiple nodes, availability zones, regions, or infrastructure domains. If one instance fails, another replica can potentially continue serving requests. However, replication introduces additional challenges around consistency, synchronisation, failover, and data correctness. Simply creating multiple copies does not automatically create a reliable system.

Quorum and Distributed Consensus:

When multiple replicas maintain shared state, the system needs rules for deciding which operations or versions of data should be accepted. Quorum-based approaches require a sufficient number of replicas to participate in an operation before it is considered successful. Distributed consensus algorithms such as Raft and Paxos provide mechanisms for multiple nodes to agree on a sequence of state changes despite certain failures. These mechanisms are particularly important for distributed databases, coordination services, and control planes.

Consistency and Availability:

Distributed systems often have to make explicit trade-offs between consistency, latency, availability, and partition tolerance. Network partitions are especially important because two parts of a distributed system may temporarily be unable to communicate while both remain operational. The system must define how it behaves during this period—for example, whether it prioritises serving requests, preserving strong consistency, or temporarily rejecting operations to protect data integrity.

Timeouts and Failure Detection:

A distributed system cannot normally distinguish perfectly between a failed component and a slow component. A request that has not produced a response may be delayed, dropped, or processed by a service that is temporarily overloaded. Timeouts therefore become an important reliability mechanism. Appropriate timeout policies prevent resources from being consumed indefinitely while allowing the system to detect unhealthy dependencies and initiate recovery actions.

Retries and Backoff:

Retries can improve resilience when failures are temporary, such as transient network errors or overloaded dependencies. However, uncontrolled retries can make an outage significantly worse by generating additional traffic against an already struggling service. Reliable systems therefore commonly use bounded retries, exponential backoff, jitter, and carefully selected retry policies. Applications should also understand whether an operation is safe to repeat.

Idempotency:

Idempotency is particularly important for distributed applications that perform operations such as payments, provisioning, or database updates. If a client sends a request and does not receive a response, it may not know whether the operation succeeded. Retrying the operation could therefore create a duplicate action unless the system provides an idempotency mechanism or otherwise guarantees safe repetition.

Circuit Breakers and Load Shedding:

A service should not continue sending requests to a dependency that is consistently failing or timing out. Circuit-breaker patterns can temporarily stop calls to an unhealthy dependency and allow the system to recover without continuously generating additional load. Load shedding provides another protection mechanism by deliberately rejecting lower-priority work when resources become constrained, preserving capacity for critical operations.

Graceful Degradation:

Reliability does not always mean providing every feature under every condition. A well-designed distributed application can degrade gracefully when dependencies fail. For example, a recommendation service might become unavailable while the core shopping experience continues operating. Cached data, reduced functionality, asynchronous processing, and prioritised workloads can allow critical user journeys to remain available during partial failures.

Data Durability and Recovery:

Availability is only one dimension of reliability. A system can remain online while still losing or corrupting data. Reliable architectures therefore require durable storage, backups, replication, recovery procedures, and clearly defined recovery objectives. Recovery Point Objectives determine how much data loss can be tolerated, while Recovery Time Objectives define how quickly a service should be restored after a major failure.

Observability and Reliability Engineering:

Reliability depends on knowing what the system is actually doing. Metrics, logs, traces, events, health checks, and distributed tracing provide the telemetry required to detect failures and understand their propagation across service dependencies. Service Level Indicators and Service Level Objectives can then turn reliability into measurable engineering targets. This connects distributed-systems design with modern SRE practices and error-budget management.

Failure Isolation:

One of the most important reliability principles is preventing a local failure from becoming a system-wide failure. Resource limits, bulkheads, independent failure domains, workload isolation, queue boundaries, and carefully designed service dependencies can limit the blast radius of an incident. The objective is not to eliminate every failure, but to ensure that an individual failure does not unnecessarily cascade through the entire platform.

Testing Distributed Failure:

Reliability cannot be demonstrated simply by testing normal operation. Systems should also be tested under realistic failure conditions. Controlled experiments can introduce network latency, packet loss, service failures, node termination, resource exhaustion, dependency failures, and other faults. Techniques such as chaos engineering can help determine whether the architecture actually behaves as expected when components fail.

Designing for Recovery:

Recovery should be treated as part of the architecture rather than as an emergency procedure added later. Services should have clearly defined health checks, restart behaviour, failover mechanisms, backup strategies, recovery procedures, and operational runbooks. Automated recovery can reduce the duration of incidents, but automation must also be designed carefully so that recovery mechanisms do not amplify an existing failure.

The key principle behind reliable distributed systems is simple: assume that components will fail, communication will sometimes be unreliable, and unexpected states will occur. Reliability comes from designing the system so that these failures are detected, isolated, contained, and recovered from without unnecessarily affecting users or compromising data.

Modern distributed systems therefore combine multiple reliability mechanisms rather than relying on a single technology. Replication provides redundancy, consensus coordinates shared state, timeouts and retries manage unreliable communication, circuit breakers and load shedding contain failures, observability provides operational visibility, and recovery mechanisms restore service after serious incidents.

Ultimately, reliability is an architectural property as much as an operational one. A distributed system becomes resilient when failure is treated as a normal condition that the design is explicitly prepared to handle. By combining fault isolation, replication, consistency strategies, graceful degradation, observability, and tested recovery procedures, organisations can build platforms that continue delivering critical services even when individual components—and sometimes entire infrastructure domains—fail.

Interactive Distributed Systems Lab 01

Distributed Reliability Decision Lab

Reliability engineering is about choosing the right mechanism for the failure mode. Diagnose the scenario, identify the dominant risk, and select the reliability pattern that best addresses it.

ENGINEERING CHALLENGE

A distributed application is experiencing intermittent failures. The system contains multiple services, replicated workloads and external dependencies. Your job is to choose the reliability mechanism that addresses each failure without simply adding more capacity.

Failure scenario
Scenario 1

Choose the reliability mechanism
Reliability assessment complete.
You have worked through five different failure modes. The key lesson is that reliability is not one feature or one product. Redundancy, replication, retries, circuit breaking, bulkheads and graceful degradation solve different classes of failure.
Engineering takeaway: Do not treat every distributed-system failure as a capacity problem. First identify the failure mode, then apply the reliability mechanism that limits its blast radius, preserves correctness, or enables recovery.

Highlights: Reliability In Distributed System

Understanding Distributed Systems

At their core, distributed systems consist of multiple interconnected nodes working together to achieve a common goal. These nodes can be geographically dispersed and communicate through various protocols. Understanding the structure and behavior of distributed systems is crucial before exploring reliability measures.

To grasp the inner workings of distributed systems, it’s essential to familiarize ourselves with their key components. These include communication protocols, consensus algorithms, fault tolerance mechanisms, and distributed data storage. Each component plays a crucial role in ensuring the reliability and efficiency of distributed systems.

– Challenges and Risks: Reliability in distributed systems faces several challenges and risks due to their inherent nature. Network failures, node crashes, message delays, and data inconsistency are common issues compromising system reliability. Furthermore, the complexity of these systems amplifies the difficulty of diagnosing and resolving failures promptly.

– Replication and Redundancy: To mitigate the risks associated with distributed systems, replication and redundancy techniques are employed. Replicating data and functionalities across multiple nodes ensures fault tolerance and enhances reliability. The system can continue to operate with redundant components even if specific nodes fail.

– Consistency and Coordination: Maintaining data consistency is crucial in distributed systems. Distributed consensus protocols, such as the Paxos or Raft consensus algorithms, ensure that all nodes agree on the same state despite failures or network partitions. Coordinating actions among distributed nodes is essential to prevent conflicts and ensure reliable system behavior.

– Monitoring and Failure Detection: Continuous monitoring and failure detection mechanisms are essential for identifying and resolving issues promptly. Various monitoring tools and techniques, such as heartbeat protocols and health checks, can help detect failures and initiate recovery processes. Proactive monitoring and regular maintenance significantly contribute to the overall reliability of distributed systems.

**Shift in Landscape**

When considering reliability in a distributed system, considerable shifts in our environmental landscape have caused us to examine how we operate and run our systems and networks. We have had a mega change with the introduction of various cloud platforms and their services and containers.

In addition, with the complexity of managing distributed systems observability and microservices observability that unveil significant gaps in current practices in our technologies. Not to mention the flaws with the operational practices around these technologies.

Reliability in Distributed Systems

Reliability in distributed systems means ensuring that a system continues to operate correctly even when individual components fail, networks partition, nodes crash, messages arrive late, or workloads spike unexpectedly. Distributed systems are inherently unreliable — so reliability must be engineered deliberately through redundancy, consistency models, fault‑tolerance, and observability.

Below is a structured, high‑clarity section you can drop directly into your distributed systems or SRE content.

Core Goal of Reliability

A reliable distributed system continues to deliver correct results despite failures, despite scale, and despite complexity.

Reliability is achieved through:

  • fault tolerance

  • redundancy

  • consistency guarantees

  • replication

  • observability

  • SLOs

  • graceful degradation

Why Reliability Is Hard in Distributed Systems

Distributed systems fail in non‑linear ways:

  • partial failures

  • network partitions

  • retry storms

  • cascading latency

  • inconsistent replicas

  • clock drift

  • leader election failures

  • multi‑region divergence

Reliability must be engineered at every layer.

 

Key Reliability Techniques

1. Replication

Replicate data and services across nodes or regions:

  • active‑active

  • active‑standby

  • quorum‑based replication

  • multi‑region replication

Ensures availability even if nodes fail.

2. Fault Tolerance

Systems must continue operating during failures:

  • retry logic

  • circuit breakers

  • backpressure

  • failover routing

  • leader election (Raft, Paxos)

Fault tolerance prevents cascading failures.

3. Consistency Models

Choose the right consistency level:

  • strong consistency

  • eventual consistency

  • causal consistency

  • read‑your‑writes

Consistency affects correctness and user experience.

4. Redundancy

Redundant:

  • nodes

  • services

  • network paths

  • storage

  • load balancers

Redundancy eliminates single points of failure.

5. Graceful Degradation

If parts fail, the system still works:

  • disable non‑critical features

  • reduce load

  • serve cached data

  • fallback logic

Users experience reduced functionality, not outages.

6. Observability

Reliability requires visibility:

  • metrics

  • logs

  • traces

  • dependency maps

  • SLO burn rate

Observability reveals why failures occur.

Reliability Patterns in Distributed Systems

Circuit Breaker Pattern

Stops calling failing services to prevent cascading failures.

Bulkhead Pattern

Isolates components so one failure doesn’t sink the entire system.

Retry + Backoff

Retries failed operations with exponential backoff to avoid overload.

Leader Election

Ensures a single authoritative node manages coordination.

Idempotency

Allows safe retries without duplicating operations.

Event Sourcing

Stores state as events for replay and recovery.

Chaos Engineering

Proactively tests reliability by injecting controlled failures.

 

Reliability in Cloud & Kubernetes

Distributed systems today run on:

  • Kubernetes

  • microservices

  • serverless

  • multi‑region cloud

Reliability requires:

  • HPA/VPA auto‑scaling

  • pod disruption budgets

  • node health checks

  • multi‑AZ deployments

  • service mesh retries + mTLS

  • distributed tracing

Outcome

Reliable distributed systems deliver:

  • predictable performance

  • resilience to node and network failures

  • consistent data across replicas

  • fast recovery from outages

  • strong SLO compliance

  • reduced downtime

  • improved user experience

Reliability transforms distributed systems from fragile, complex architectures into robust, fault‑tolerant, self‑healing platforms.

Managed Instance Groups (MIGs)

**Understanding the Basics of Managed Instance Groups**

Managed Instance Groups are collections of virtual machine (VM) instances that are treated as a single entity. They are designed to simplify the management of multiple instances by automating tasks like scaling, updating, and load balancing. With MIGs, you can ensure that your application has the right number of instances running at any given time, responding dynamically to changes in demand.

Google Cloud’s MIGs make it easy to deploy applications with high availability and reliability. By using templates, you can define the configuration for all instances in the group, ensuring consistency and reducing the potential for human error.

—

**Ensuring Reliability in Distributed Systems**

Reliability is a critical component of any distributed system, and Managed Instance Groups play a significant role in achieving it. By distributing workloads across multiple instances and regions, MIGs help prevent single points of failure. If an instance fails, the group automatically replaces it with a new one, minimizing downtime and ensuring continuous service availability.

Moreover, Google Cloud’s infrastructure ensures that your instances are backed by a robust network, providing low-latency access and high-speed connectivity. This further enhances the reliability of your applications and services, giving you peace of mind as you scale.

—

**Scaling with Ease and Flexibility**

One of the standout features of Managed Instance Groups is their ability to scale quickly and efficiently. Whether you’re dealing with sudden spikes in traffic or planning for steady growth, MIGs offer flexible scaling policies to meet your needs. You can scale based on CPU utilization, load balancing capacity, or even custom metrics, allowing for precise control over your application’s performance.

Google Cloud’s autoscaling capabilities mean you only pay for the resources you use, optimizing cost-efficiency while maintaining high performance. This flexibility makes MIGs an ideal choice for businesses looking to grow their cloud infrastructure without unnecessary expenditure.

—

**Integrating with Google Cloud Ecosystem**

Managed Instance Groups seamlessly integrate with other Google Cloud services, providing a cohesive ecosystem for your applications. They work in harmony with Cloud Load Balancing to distribute traffic efficiently and with Stackdriver for monitoring and logging, giving you comprehensive insights into your application’s performance.

By leveraging Google Cloud’s extensive suite of tools, you can build, deploy, and manage applications with greater agility and confidence. This integration streamlines operations and simplifies the complexities of managing a distributed system.

Managed Instance Group

Example Product: Cisco AppDynamics

### What is Cisco AppDynamics?

Cisco AppDynamics is an application performance management (APM) solution that offers deep insights into your application’s performance, user experience, and business impact. It helps IT teams detect, diagnose, and resolve issues quickly, ensuring a seamless digital experience for end-users. By leveraging machine learning and artificial intelligence, AppDynamics provides actionable insights to optimize your applications.

### Key Features of Cisco AppDynamics

#### Real-Time Performance Monitoring

With Cisco AppDynamics, you can monitor the performance of your applications in real-time. This feature allows you to detect anomalies and performance issues as they happen, ensuring that you can address them before they impact your users.

#### End-User Monitoring

Understanding how your users interact with your applications is crucial. AppDynamics offers end-user monitoring, which provides visibility into the user journey, from the front-end user interface to the back-end services. This helps you identify and resolve issues that directly affect user experience.

#### Business Transaction Monitoring

AppDynamics breaks down your application into business transactions, which are critical user interactions within your application. By monitoring these transactions, you can gain insights into how your application supports key business processes and identify areas for improvement.

#### AI-Powered Analytics

The platform’s AI-powered analytics enable you to predict and prevent performance issues before they occur. By analyzing historical data and identifying patterns, AppDynamics helps you proactively manage your application’s performance.

### Benefits of Using Cisco AppDynamics

#### Improved Application Performance

By continuously monitoring your application’s performance, AppDynamics helps you identify and resolve issues quickly, ensuring that your application runs smoothly and efficiently.

#### Enhanced User Experience

With end-user monitoring, you can gain insights into how users interact with your application and address any issues that may affect their experience. This leads to increased user satisfaction and retention.

#### Better Business Insights

Business transaction monitoring provides a clear understanding of how your application supports critical business processes. This helps you make data-driven decisions to optimize your application and drive business growth.

Monitoring GKE Environment

The Significance of Monitoring in GKE

Monitoring in GKE goes beyond simply monitoring resource utilization. It provides valuable insights into the health and performance of your Kubernetes clusters, nodes, and applications running within them. You gain a comprehensive understanding of your system’s behavior by closely monitoring key metrics such as CPU usage, memory utilization, and network traffic. You can proactively address issues before they escalate.

GKE-Native Monitoring has many powerful features that simplify the monitoring process. One notable feature is integrating with Stackdriver, Google Cloud’s monitoring and observability platform.

With this integration, you can access a rich set of monitoring tools, including customizable dashboards, alerts, and logging capabilities, all designed explicitly for GKE deployments. Additionally, GKE-Native Monitoring seamlessly integrates with other Google Cloud services, enabling you to leverage advanced analytics and machine learning capabilities.

Challenges to Gaining Reliability

**Existing Static Tools**

This has caused a knee-jerk reaction to a welcomed drive-in innovation to system reliability. Yet, some technologies and tools used to manage these innovations do not align with the innovative events. Many of these tools have stayed relatively static in our dynamic environment. So, we have static tools used in a dynamic environment, which causes friction to reliability in distributed systems and the rise for more efficient network visibility.

**Understanding the Complexity**

Distributed systems are inherently complex, with multiple components across different machines or networks. This complexity introduces challenges like network latency, hardware failures, and communication bottlenecks. Understanding the intricate nature of distributed systems is crucial to devising reliable solutions.

Gaining Reliability

**Redundancy and Replication**

One critical approach to enhancing reliability in distributed systems is redundancy and replication. By duplicating critical components or data across multiple nodes, the system becomes more fault-tolerant. This ensures the system can function seamlessly even if one component fails, minimizing the risk of complete failure.

**Consistency and Consensus Algorithms**

Maintaining consistency in distributed systems is a significant challenge due to the possibility of concurrent updates and network delays. Consensus algorithms, such as the Paxos or Raft algorithms, are vital in achieving consistency by ensuring agreement among distributed nodes. These algorithms enable reliable decision-making and guarantee that all nodes reach a consensus state.

**Monitoring and Failure Detection**

To ensure reliability, robust monitoring mechanisms are essential. Monitoring tools can track system performance, resource utilization, and network health. Additionally, implementing efficient failure detection mechanisms allows for prompt identification of faulty components, enabling proactive measures to mitigate their impact on the overall system.

**Load Balancing and Scalability**

Load balancing is crucial in distributing the workload evenly across nodes in a distributed system. It ensures that no single node is overwhelmed, reducing the risk of system instability. Furthermore, designing systems with scalability in mind allows for seamless expansion as the workload grows, ensuring that reliability is maintained even during periods of high demand.

Required: Distributed Tracing

Using distributed tracing, you can profile or monitor the results of requests across a distributed system. Distributed systems can be challenging to monitor since each node generates its logs and metrics. To get a complete view of a distributed system, it is necessary to aggregate these separate node metrics holistically. 

A distributed system generally doesn’t access its entire set of nodes but rather a path through those nodes. With distributed tracing, teams can analyze and monitor commonly accessed paths through a distributed system. The distributed tracing is installed on each system node, allowing teams to query the system for information on node health and performance.

Benefits: Distributed Tracing

Despite the challenges, distributed systems offer a wide array of benefits. One notable advantage is enhanced fault tolerance. Distributing tasks and data across multiple nodes improves system reliability, as a single point of failure does not bring down the entire system.

Additionally, distributed systems enable improved scalability, accommodating growing demands by adding more nodes to the network. The applications of distributed systems are vast, ranging from cloud computing and large-scale data processing to peer-to-peer networks and distributed databases.

Google Cloud Trace

Understanding Cloud Trace

Cloud Trace, an integral part of Google Cloud’s observability offerings, provides developers with a detailed view of their application’s performance. It enables tracing and analysis of requests as they flow through different components of a distributed system. By visualizing the latency, bottlenecks, and dependencies, Cloud Trace empowers developers to optimize their applications for better performance and user experience.

Cloud Trace offers a range of features to simplify the monitoring and troubleshooting process. With its distributed tracing capabilities, developers can gain insights into how requests traverse various services, identify latency issues, and pinpoint the root causes of performance bottlenecks. The integration with Google Cloud’s ecosystem allows seamless correlation between traces and other monitoring data, enabling a comprehensive view of application health.

Improve Performance & Reliability 

By leveraging Cloud Trace, developers can significantly improve the performance and reliability of their applications. The ability to pinpoint and resolve performance issues quickly translates into enhanced user satisfaction and higher productivity. Moreover, Cloud Trace enables proactive monitoring, ensuring that potential bottlenecks and inefficiencies are identified before they impact end-users. For organizations, this translates into cost savings, improved scalability, and better resource utilization.

Adopting Cloud Trace

Cloud Trace has been adopted by numerous organizations across various industries, with remarkable outcomes. From optimizing the response time of e-commerce platforms to enhancing the efficiency of complex microservices architectures, Cloud Trace has proven its worth in diagnosing performance issues and driving continuous improvement. The real-time visibility provided by Cloud Trace empowers organizations to make data-driven decisions and deliver exceptional user experiences.

 

Interactive Distributed Systems Lab 02

Failure & Recovery Simulator

Distributed systems rarely fail cleanly. Inject a fault, observe how the failure propagates, detect the problem, and then recover the affected component.

SIMULATION OBJECTIVE

Select a failure mode and follow the recovery lifecycle: Healthy → Failure → Detection → Recovery. Notice how a local fault can create a wider degraded state when the architecture does not isolate the failure.

Live service topology
⇄
Load Balancer
Healthy
▣
Service A
Healthy
◈
Service B
Healthy
▤
Database
Healthy
Inject a failure
System state
Overall service HEALTHY
Failure detection READY
Recovery STANDBY
Event console

Reliability In Distributed System

Adopting Distributed Systems

Distributed systems refer to a network of interconnected computers that communicate and coordinate their actions to achieve a common goal. Unlike traditional centralized systems, where a single entity controls all components, distributed systems distribute tasks and data across multiple nodes. This decentralized approach enables enhanced scalability, fault tolerance, and resource utilization.

**Key Components of Distributed Systems**

To comprehend the inner workings of distributed systems, we must familiarize ourselves with their key components. These components include nodes, communication channels, protocols, and distributed file systems. Nodes represent individual machines or devices within the network; communication channels facilitate data transmission, protocols ensure reliable communication and distributed file systems enable data storage across multiple nodes.

Distribued vs centralized

**Distributed Systems Use Cases**

Many modern applications use distributed systems, including mobile and web applications with high traffic. Web browsers or mobile applications serve as clients in a client-server environment, and the server becomes its own distributed system. The modern web server follows a multi-tier system pattern. Requests are delegated to several server logic nodes via a load balancer.

Kubernetes is popular among distributed systems since it enables containers to be combined into a distributed system. Kubernetes orchestrates network communication between the distributed system nodes and handles dynamic horizontal and vertical scaling of the nodes. 

Cryptocurrencies like Bitcoin and Ethereum are also peer-to-peer distributed systems. The currency ledger is replicated at every node in a cryptocurrency network. To bootstrap, a currency node connects to other nodes and downloads its full ledger copy. Additionally, cryptocurrency wallets use JSON RPC to communicate with the ledger nodes.

Challenges in Distributed Systems

While distributed systems offer numerous advantages, they also pose various challenges. One significant challenge is achieving consensus among distributed nodes. Ensuring that all nodes agree on a particular value or decision can be complex, especially in the presence of failures or network partitions. Additionally, maintaining data consistency across distributed nodes and mitigating issues related to concurrency control requires careful design and implementation.

**Example: Distributed System of Microservices**

Microservices are one type of distributed system since they decompose an application into individual components. A microservice architecture, for example, may have services corresponding to business features (payments, users, products, etc.), with each element handling the corresponding business logic. Multiple redundant copies of the services will then be available, so there is no single point of failure.

**Distributed Systems: The Challenge**

Distributed systems are required to implement the reliability, agility, and scale expected of modern computer programs. Distributed systems are applications of many different components running on many other machines. Containers are the foundational building block, and groups of containers co-located on a single device comprise the atomic elements of distributed system patterns.

The significant shift we see with software platforms is that they evolve much quicker than the products and paradigms we use to monitor them. We need to consider new practices and technologies with dedicated platform teams to enable a new era of system reliability in a distributed system. Along with the practices of Observability that are a step up to the traditional monitoring of static infrastructure: Observability vs monitoring.

Knowledge Check: Distributed Systems Architecture

  • Client-Server Architecture

A client-server architecture has two primary responsibilities. The client presents user interfaces and is connected to the server via a network. The server handles business logic and state management. Unless the server is redundant, a client-server architecture can quickly degrade into a centralized architecture. A truly distributed client-server setup will consist of multiple server nodes that distribute client connections. In modern client-server architectures, clients connect to encapsulated distributed systems on the server.

  • Multi-tier Architecture

Multi-tier architectures are extensions of client-server architectures. Multi-tier architectures decompose servers into further granular nodes, decoupling additional backend server responsibilities like data processing and management. By processing long-running jobs asynchronously, these additional nodes free up the remaining backend nodes to focus on responding to client requests and interacting with the data store.

  • Peer-to-Peer Architecture

Peer-to-peer distributed systems contain complete instances of applications on each node. There is no separation between presentation and data processing at the node level. A node consists of a presentation layer and a data handling layer. Peer nodes may contain the entire system’s state data. 

Peer-to-peer systems have a great deal of redundancy. When initiated and brought online, peer-to-peer nodes discover and connect to other peers, thereby synchronizing their local state with the system’s. As a result of this feature, nodes on a peer-to-peer network won’t be disrupted by the failure of one. Additionally, peer-to-peer systems will persist. 

  • Service-orientated Architecture

A service-oriented architecture (SOA) is a precursor to microservices. Microservices differ from SOA primarily in their node scope, which is at the feature level. Each microservice node encapsulates a specific set of business logic, such as payment processing—multiple nodes of business logic interface with independent databases in a microservice architecture. In contrast, SOA nodes encapsulate an entire application or enterprise division. Database systems are typically included within the service boundary of SOA nodes.

Because of their benefits, microservices have become more popular than SOA. The small service nodes provide functionality that teams can reuse through microservices. The advantages of microservices include greater robustness and a more extraordinary ability for vertical and horizontal scaling to be dynamic.

Reliability in Distributed Systems: Components

A- Redundancy and Replication:

Redundancy and replication are two fundamental concepts distributed systems use to enhance reliability. Redundancy involves duplicating critical system components, such as servers, storage devices, or network links, so the redundant component can seamlessly take over if one fails. Replication, on the other hand, involves creating multiple copies of data across different nodes in a system, enabling efficient data access and fault tolerance. By incorporating redundancy and replication, distributed systems can continue to operate even when individual components fail.

B – Fault Tolerance:

Fault tolerance is a crucial aspect of achieving reliability in distributed systems. It involves designing systems to operate correctly even when one or more components encounter failures. Several techniques, such as error detection, recovery, and prevention mechanisms, are employed to achieve fault tolerance.

C – Error Detection:

Error detection techniques, such as checksums, hashing, and cyclic redundancy checks (CRC), identify errors or data corruption during transmission or storage. By verifying data integrity, these techniques help identify and mitigate potential failures in distributed systems.

D – Error Recovery:

Error recovery mechanisms, such as checkpointing and rollback recovery, aim to restore the system to a consistent state after a failure. Checkpointing involves periodically saving the system’s state and data, allowing recovery to a previously known good state in case of failures. On the other hand, rollback recovery involves undoing the effects of failed operations and returning the system to a consistent state.

E – Error Prevention:

Distributed systems employ error prevention techniques, such as redundancy elimination, consensus algorithms, and load balancing to enhance reliability. Redundancy elimination reduces unnecessary duplication of data or computation, thereby reducing the chances of errors. Consensus algorithms ensure that all nodes in a distributed system agree on a shared state despite failures or message delays. Load balancing techniques distribute computational tasks evenly across multiple nodes to prevent overloading and potential shortcomings.

Challenges: Traditional Monitoring

**Lack of Connective Event**

If you examine traditional monitoring systems, they look to capture and investigate signals in isolation. They work in a siloed environment, similar to that of developers and operators before the rise of DevOps. Existing monitoring systems cannot detect the “Unknowns Unknowns” that are familiar with modern distributed systems. This often leads to service disruptions. So, you may be asking what an “Unknown Unknown” is.

I’ll put it to you this way: the distributed systems we see today lack predictability—certainly not enough predictability to rely on static thresholds, alerts, and old monitoring tools. If something is fixed, it can be automated, and we have static events, such as in Kubernetes, a POD reaching a limit.

Then, a replica set introduces another pod on a different node if specific parameters are met, such as Kubernetes Labels and Node Selectors. However, this is only a tiny piece of the failure puzzle in a distributed environment.  Today, we have what’s known as partial failures and systems that fail in very creative ways.

Reliability In Distributed System: Creative ways to fail

So, we know that some of these failures are quickly predicted, and actions are taken. For example, if this Kubernetes POD node reaches a specific utilization, we can automatically reschedule PODs on a different node to stay within our known scale limits.

Predictable failures can be automated in Kubernetes and with any infrastructure. An Ansible script is useful when these events occur. However, we have much more to deal with than POD scaling; we have many partial and complicated failures known as black holes.

**In today’s world of partial failures**

Microservices applications are distributed and susceptible to many external factors. On the other hand, if you examine the traditional monolithic application style, all the functions reside in the same process. It was either switched ON or OFF!! Not much happened in between. So, if there is a failure in the procedure, the application as a whole will fail. The results are binary, usually either a UP or Down.

This was easy to detect with some essential monitoring, and failures were predictable. There was no such thing as a partial failure. In a monolith application, all application functions are within the same process. A significant benefit of these monoliths is that you don’t have partial failures.

However, in a cloud-native world, where we have broken the old monolith into a microservices-based application, a client request can go through multiple hops of microservices, and we can have several problems to deal with.

There is a lack of connectivity between the different domains. Many monitoring tools and knowledge will be tied to each domain, and alerts are often tied to thresholds or rate-of-change violations that have nothing to do with user satisfaction, which is a critical metric to care about.

**System reliability: Today, you have no way to predict**

So, the new, modern, and complex distributed systems place very different demands on your infrastructure—considerably different from the simple three-tier application, where everything is generally housed in one location.  We can’t predict anything anymore, which breaks traditional monitoring approaches.

When you can no longer predict what will happen, you can no longer rely on a reactive approach to monitoring and management. The move towards a proactive approach to system reliability is a welcomed strategy.

**Blackholes: Strange failure modes**

When considering a distributed system, many things can happen. A service or region can disappear or disappear for a few seconds or ms and reappear. We believe this is going into a black hole when we have strange failure modes. So when anything goes into it will disappear. Peculiar failure modes are unexpected and surprising.

Strange failure modes are undoubtedly unpredictable. So, what happens when your banking transactions are in a black hole? What if your banking balance is displayed incorrectly or if you make a transfer to an external account and it does not show up? 

Site Reliability Engineering (SRE) and Observability

Site reliability engineering (SRE) and observational practices are needed to manage these types of unpredictability and unknown failures. SRE is about making systems more reliable. And everyone has a different way of implementing SRE practices. Usually, about 20% of your issues cause 80% of your problems.

You need to be proactive and fix these issues upfront. You need to be able to get ahead of the curve and do these things to prevent incidents from occurring. This usually happens in the wake of a massive incident. This usually acts as a teachable moment. It gives the power to be the reason to listen to a Chaos Engineering project. 

New tools and technologies:

1 – Distributed tracing

We have new tools, such as distributed tracing. So, what is the best way to find the bottleneck if the system becomes slow? Here, you can use Distributed Tracing and Open Telemetry. The tracing helps us instrument our system, figuring out where the time has been spent and where it can be used across distributed microservice architecture to troubleshoot problems. Open Telemetry provides a standardized way of instrumenting our system and providing those traces.

2 – SLA, SLI, SLO, and Error Budgets

So we don’t just want to know when something has happened and then react to an event that is not looking from the customer’s perspective. We need to understand if we are meeting SLA by gathering the number and frequency of the outages and any performance issues.

Service Level Objectives (SLO) and Service Level Indicators (SLI) can assist you with measurements. Service Level Objectives (SLOs) and Service Level Indicators (SLI) not only help you with measurements but also offer a tool for having better reliability and forming the base for the reliability stack.

Interactive Distributed Systems Lab 03

Consistency & Consensus Lab

Replication improves availability and durability, but multiple copies create coordination problems. Explore replica lag, leader failure, network partitions and quorum decisions while keeping consistency and consensus concepts separate.

LAB OBJECTIVE

Use the controls to create distributed-system events. Observe what happens to replica state, then determine whether the system can safely continue accepting writes or needs a new leader and quorum.

Replicated data service
◆
Replica 1
Leader
v104
◇
Replica 2
Follower
v104
◇
Replica 3
Follower
v104
Inject an event
Quorum: In this three-replica example, a majority is two replicas. A consensus protocol can use quorum rules to prevent conflicting leaders or committed decisions.
Distributed-system analysis
System ready.
All three replicas contain the same committed version. Select an event to explore what changes.
Property
Current
Meaning
Leader
R1
Active
Quorum
3 / 3
Available
Versions
104 / 104 / 104
Aligned
Chaos Engineering

Baseline Engineering

Baseline Engineering

In modern networks, “normal” is not a single number. Network behaviour changes throughout the day as users connect, applications scale, workloads move between data centres and cloud platforms, and traffic patterns shift. Network baseline engineering provides a structured way to understand that normal behaviour and use it as a reference for detecting meaningful changes.

A network baseline is a measured representation of expected network behaviour over time. Instead of relying on a single bandwidth or latency value, an effective baseline can include interface utilisation, packet loss, latency, jitter, application response time, routing changes, flow volumes, errors and discards, CPU and memory utilisation, connection rates, and other service-specific indicators.

The objective is not simply to record historical statistics. A useful baseline provides operational context: is the current behaviour expected, unusual, or indicative of a developing problem?

Proactive Anomaly Detection

Once normal behaviour has been established, current telemetry can be compared against the expected range. A sudden increase in packet loss, an unusual latency pattern, unexpected interface utilisation, or a significant change in traffic flows can indicate congestion, equipment failure, routing problems, security events, or application issues.

Modern monitoring platforms can go beyond simple threshold alerts by using historical trends, statistical ranges, seasonal patterns, and anomaly detection. This helps reduce the number of alerts generated by normal variations while highlighting changes that deserve investigation.

Performance Optimisation

A baseline also provides the evidence needed to optimise network performance. Engineers can compare utilisation and application behaviour across different periods, identify persistent bottlenecks, and determine whether performance problems are caused by network capacity, routing, infrastructure, or dependent services.

For example, consistently high utilisation on an uplink during a predictable business period may represent normal demand rather than an incident. However, if utilisation suddenly increases outside that normal pattern while latency and packet loss also rise, the baseline provides valuable context for identifying the change.

Capacity Planning

Baseline data is particularly valuable for network capacity planning. Historical utilisation trends can reveal how quickly links, devices, wireless infrastructure, VPN gateways, or cloud connectivity are approaching their practical limits.

Capacity decisions can therefore be based on observed demand rather than assumptions. Engineers can identify sustained growth, recurring peak periods, and infrastructure approaching saturation before performance degradation becomes a user-facing problem.

Building a Network Baseline

The first stage is collecting representative telemetry. Depending on the environment, this may include interface counters, SNMP metrics, streaming telemetry, NetFlow or IPFIX records, application performance measurements, routing information, device health statistics, logs, and synthetic tests.

The collection period should cover enough normal operating conditions to reveal meaningful patterns. A business network, for example, may behave very differently during working hours, evenings, weekends, maintenance windows, and scheduled batch-processing periods.

Baseline Analysis

Collected data can then be analysed to identify typical behaviour and acceptable variation. Useful statistical measures can include averages, medians, percentiles, standard deviation, minimum and maximum values, and historical ranges.

Percentiles are particularly useful for network engineering because averages can hide short periods of severe degradation. For latency-sensitive services, for example, the 95th or 99th percentile may provide a much better representation of user experience than the average latency alone.

Baselines should also be segmented where appropriate. A single baseline for an entire network may be too broad to be useful. Separate baselines can be created for different interfaces, sites, applications, services, traffic classes, time periods, or business-critical workloads.

Baseline Maintenance

A network baseline is not a static document. Infrastructure changes, application deployments, cloud migrations, new users, routing changes, and changes in business activity can all alter normal network behaviour.

The baseline should therefore evolve with the environment. However, changes should not automatically be incorporated into the baseline without investigation. Otherwise, a gradual performance problem could eventually become classified as “normal.”

A mature approach separates genuine changes in expected behaviour from degradation that should trigger engineering investigation.

From Monitoring to Network Intelligence

Network baseline engineering becomes significantly more powerful when combined with modern observability. Metrics can show what changed, flow telemetry can help explain where traffic moved, logs can provide operational context, and traces or application telemetry can help determine whether a network symptom is actually caused by an application or dependency.

This creates a more complete operational workflow:

Measure → Establish Normal Behaviour → Detect Deviation → Investigate Context → Remediate → Reassess the Baseline

The goal is not to eliminate variation from the network. Networks are dynamic systems and variation is expected. The goal is to understand that variation well enough to distinguish normal behaviour from meaningful change.

A well-engineered network baseline therefore becomes more than a historical performance report. It provides an operational reference for troubleshooting, anomaly detection, performance optimisation, capacity planning, and long-term infrastructure decisions.

For modern networks, establishing what “normal” looks like is one of the most effective foundations for recognising when something has genuinely changed.

Interactive Baseline Engineering Lab 01

Network Baseline Builder

Establish a measurable picture of normal network behaviour, then compare a later live sample against it. The goal is to turn “something feels slow” into evidence-based baseline drift.

Scenario

You are establishing a baseline for a distributed application during normal business operation. Record representative values first. Then test a later sample to determine whether the environment remains within its expected range.

1. Normal Operating Sample

Known-good state
WAN Latency ms
22 ms
Jitter ms
4 ms
Packet Loss %
0.2 %
Throughput Mbps
640 Mbps
CPU Utilisation %
48 %
App Response ms
180 ms

2. Baseline Profile

Not created
Set the normal operating values on the left and select Create Baseline.

3. Live Sample

Compare current telemetry
WAN Latency ms
24 ms
Jitter ms
5 ms
Packet Loss %
0.2 %
Throughput Mbps
625 Mbps
CPU Utilisation %
52 %
App Response ms
190 ms
Baseline Comparison
WAITING
Why baseline engineering matters

A baseline is not simply a target number. It is a measured representation of normal behaviour under defined conditions. A useful baseline should be representative, repeatable and connected to operational or user expectations. Thresholds in this simulation are teaching tolerances, not universal production values.

ENGINEERING TAKEAWAY

Measure normal behaviour first. Once you have a reliable reference point, live telemetry can be compared against it to identify drift, investigate anomalies and decide whether a change requires action.

Highlights: Baseline Engineering

1. Understanding Baseline Engineering

Baseline engineering serves as the bedrock for any engineering project. It involves creating a reference point or baseline from which all measurements, evaluations, and improvements are made. By establishing this starting point, engineers gain insights into the project’s progress, performance, and potential deviations from the original plan.

Baseline engineering follows a systematic and structured approach. It starts with defining project objectives and requirements, then data collection and analysis. This data provides a snapshot of the initial conditions and helps engineers set realistic targets and benchmarks. Through careful monitoring and periodic assessments, deviations from the baseline can be detected early, enabling timely corrective actions.

2. Traditional Network Infrastructure

Baseline Engineering was easy in the past; applications ran in single private data centers, potentially two data centers for high availability. There may have been some satellite PoPs, but generally, everything was housed in a few locations. These data centers were on-premises, and all components were housed internally. As a result, troubleshooting, monitoring, and baselining any issues was relatively easy. The network and infrastructure were pretty static, the network and security perimeters were known, and there weren’t many changes to the stack, for example, daily.

3. Distributed Applications

However, nowadays, we are in a completely different environment. We have distributed applications with components/services located in many other places and types of places, on-premises and in the cloud, with dependencies on both local and remote services. We span multiple sites and accommodate multiple workload types.

In comparison to the monolith, today’s applications have many different types of entry points to the external world. All of this calls for the practice of Baseline Engineering and Chaos engineering kubernetes so you can fully understand your infrastructure and scaling issues. 

Baseline Engineering

Baseline engineering is the discipline of establishing foundational, measurable, repeatable standards for systems before scaling, optimizing, or securing them. In distributed systems, cloud platforms, Kubernetes, SD‑WAN, SASE, or microservices, a baseline is the known good state — the reference point that allows teams to detect drift, measure reliability, enforce SLOs, and maintain operational consistency.

Think of baseline engineering as the “ground truth” for performance, security, configuration, and architecture.

Concise Takeaway

Baseline engineering defines the minimum acceptable, reproducible state of a system — covering performance, security, configuration, and operational behavior — so deviations can be detected early and corrected automatically.

 

🧩 Core Components of Baseline Engineering

  • Performance Baselines — p95 latency, throughput, CPU/memory norms, queue depth.

  • Security Baselines — RBAC, network policies, encryption, patch levels.

  • Configuration Baselines — infrastructure-as-code templates, cluster configs, node settings.

  • Reliability Baselines — SLOs, error budgets, failover behavior.

  • Network Baselines — WAN latency, SD‑WAN path performance, DNS resolution times.

  • Observability Baselines — metrics, logs, traces, dashboards, alert thresholds.

These baselines define what “normal” looks like.

 

Why Baseline Engineering Matters

Distributed systems drift. Cloud environments change. Kubernetes clusters evolve. SD‑WAN paths fluctuate. Without baselines:

  • anomalies go undetected

  • performance regressions hide

  • security gaps widen

  • reliability becomes unpredictable

  • SLOs burn without warning

  • troubleshooting becomes guesswork

Baseline engineering creates predictability.

 

Baseline Engineering in Modern Architectures

1. Kubernetes & OpenShift

Baselines include:

  • pod startup time

  • node pressure thresholds

  • HPA/VPA scaling behavior

  • network overlay latency

  • API server performance

This allows early detection of cluster degradation.

2. Distributed Systems

Baselines include:

  • service‑to‑service latency

  • retry patterns

  • queue depth norms

  • database response time

  • cache hit ratios

This prevents cascading failures.

3. Cloud & Serverless

Baselines include:

  • cold‑start latency

  • concurrency limits

  • API gateway performance

  • IAM permission behavior

This ensures predictable scaling and routing.

4. SD‑WAN & SASE

Baselines include:

  • path latency

  • jitter

  • packet loss

  • PoP reachability

  • DNS resolution time

This enables fast detection of ISP or routing issues.

 

How Baseline Engineering Works (Step‑by‑Step)

  1. Define critical metrics (latency, error rate, throughput).

  2. Measure normal behavior over a stable period.

  3. Set thresholds based on SLOs and user expectations.

  4. Automate baseline enforcement using IaC, GitOps, and policy engines.

  5. Continuously compare live telemetry against baselines.

  6. Alert on drift (performance, security, configuration).

  7. Auto‑remediate deviations where possible.

  8. Review baselines regularly as systems evolve.

Baseline engineering is a living process — not a one‑time setup.

Baseline Engineering Best Practices

  • define baselines early, before scaling

  • use real user behavior to set thresholds

  • automate everything (GitOps, IaC, policy engines)

  • integrate baselines with observability pipelines

  • tie baselines to SLOs and error budgets

  • update baselines after major releases

  • enforce baselines across all environments (dev → prod)

Baselines must be measurable, enforceable, and observable.

Outcome

Baseline engineering delivers:

  • predictable performance

  • stable distributed systems

  • strong security posture

  • reliable auto‑scaling

  • fast anomaly detection

  • reduced MTTR

  • consistent cloud and Kubernetes operations

  • trustworthy SLO compliance

It transforms complex systems into controlled, measurable, and self‑correcting architectures.

Managed Instance Groups

**Introduction to Managed Instance Groups**

In the fast-evolving world of cloud computing, maintaining scalability, reliability, and efficiency is essential. Managed instance groups (MIGs) on Google Cloud offer an innovative solution to achieve these goals. Whether you’re a seasoned cloud engineer or new to the Google Cloud ecosystem, understanding MIGs can significantly enhance your infrastructure management capabilities.

—

**The Role of Managed Instance Groups in Baseline Engineering**

Baseline engineering focuses on establishing a stable foundation for software development and operations. Managed instance groups play a crucial role in this process by automating the deployment and scaling of virtual machines (VMs). By setting up MIGs, baseline engineering can achieve consistent performance, reduce manual intervention, and enhance the system’s resilience to changes in demand. This automation allows engineers to focus on optimizing and innovating rather than maintaining infrastructure.

—

**Automating Scalability and Load Balancing**

One of the standout features of managed instance groups is their ability to automate scalability and load balancing. As your application experiences varying levels of traffic, MIGs automatically adjust the number of instances to meet the demand. This capability ensures that your application remains responsive and cost-efficient, as you only use the resources you need. Additionally, Google Cloud’s load balancing solutions work seamlessly with MIGs, distributing traffic evenly across instances to maintain optimal performance.

—

**Achieving Reliability with Health Checks and Autohealing**

Reliability is a cornerstone of any cloud-based application, and managed instance groups provide robust mechanisms to ensure it. Through health checks, MIGs continuously monitor the status of your VMs. If an instance fails or becomes unhealthy, the autohealing feature kicks in, replacing it with a new, healthy instance. This proactive approach minimizes downtime and maintains service continuity, contributing to a more reliable application experience for end-users.

—

**Optimizing Costs with Managed Instance Groups**

Cost management is a critical consideration in cloud computing, and managed instance groups help optimize expenses. By automatically scaling the number of instances based on demand, MIGs eliminate the need for overprovisioning resources. This dynamic resource allocation ensures that you are only paying for the compute capacity you need. Moreover, when combined with Google Cloud’s pricing models, such as sustained use discounts and committed use contracts, MIGs allow for significant cost savings.

Managed Instance Group

Google Data Centers – Service Mesh

**What is Cloud Service Mesh?**

A cloud service mesh is a dedicated infrastructure layer designed to manage service-to-service communication within a microservices architecture. It provides a way to control how different parts of an application share data with one another. Essentially, it acts as a network of microservices that make up cloud applications and ensures that communication between services is secure, fast, and reliable.

**Benefits of Cloud Service Mesh**

1. **Enhanced Security**: One of the primary advantages of a cloud service mesh is the improved security it offers. By managing communication between services, it can enforce security policies, authenticate service requests, and encrypt data in transit. This reduces the risk of data breaches and unauthorized access.

2. **Increased Reliability**: Cloud service meshes enhance the reliability of service interactions. They provide load balancing, traffic routing, and failure recovery, ensuring that services remain available even in the face of failures. This is particularly crucial for applications that require high availability and resilience.

3. **Improved Observability**: With a cloud service mesh, engineering teams can gain greater visibility into the interactions between services. This includes monitoring performance metrics, logging, and tracing requests. Such observability helps in identifying and troubleshooting issues more efficiently, leading to faster resolution times.

**Impact on Baseline Engineering**

Baseline engineering involves establishing a standard level of performance, security, and reliability for an organization’s infrastructure. The introduction of a cloud service mesh has significantly impacted this field by providing a more robust foundation for managing microservices. Here’s how:

1. **Standardization**: Cloud service meshes help in standardizing the way services communicate, making it easier to maintain consistent performance across the board. This is especially important in complex systems with numerous interdependent services.

2. **Automation**: Many cloud service meshes come with built-in automation capabilities, such as automatic retries, circuit breaking, and service discovery. This reduces the manual effort required to manage service interactions, allowing engineers to focus on more strategic tasks.

3. **Scalability**: By managing service communication more effectively, cloud service meshes enable organizations to scale their applications more easily. This is crucial for businesses that experience varying levels of demand and need to adjust their resources accordingly.

Example Product: Cisco AppDynamics

### Key Features of Cisco AppDynamics

**1. Real-Time Monitoring and Analytics**

One of the standout features of Cisco AppDynamics is its ability to provide real-time monitoring and analytics. This allows businesses to gain instant visibility into application performance and user interactions. By leveraging real-time data, organizations can quickly identify performance bottlenecks, understand user behavior, and make informed decisions to enhance application efficiency.

**2. End-to-End Transaction Visibility**

With Cisco AppDynamics, you get a comprehensive view of end-to-end transactions, from user interactions to backend processes. This visibility helps in pinpointing the exact location of issues within the application stack, whether it’s in the code, database, or infrastructure. This holistic approach ensures that no problem goes unnoticed, enabling swift resolution and minimizing downtime.

**3. AI-Powered Anomaly Detection**

Cisco AppDynamics employs advanced AI and machine learning algorithms to detect anomalies in application performance. These intelligent insights help predict and prevent potential issues before they impact end users. By learning the normal behavior of your applications, the system can alert you to deviations that might signify underlying problems, allowing proactive intervention.

### Benefits of Implementing Cisco AppDynamics

**1. Enhanced User Experience**

By continuously monitoring application performance and user interactions, Cisco AppDynamics helps ensure a smooth and uninterrupted user experience. Instant alerts and detailed reports enable IT teams to address issues swiftly, reducing the likelihood of user dissatisfaction and churn.

**2. Improved Operational Efficiency**

Cisco AppDynamics automates many aspects of performance monitoring, freeing up valuable time for IT teams. This automation reduces the need for manual checks and troubleshooting, allowing teams to focus on strategic initiatives and innovation. The platform’s ability to integrate with various IT tools further streamlines operations and enhances overall efficiency.

**3. Data-Driven Decision Making**

The rich data and analytics provided by Cisco AppDynamics empower businesses to make data-driven decisions. Whether it’s optimizing application performance, planning for capacity, or enhancing security measures, the insights gained from AppDynamics drive informed strategies that align with business goals.

### Getting Started with Cisco AppDynamics

**1. Easy Deployment**

Cisco AppDynamics offers flexible deployment options, including on-premises, cloud, and hybrid environments. The straightforward installation process and intuitive user interface make it accessible even for teams with limited APM experience. Comprehensive documentation and support further ease the onboarding process.

**2. Customizable Dashboards**

Users can create customizable dashboards to monitor the metrics that matter most to their organization. These dashboards provide at-a-glance views of key performance indicators (KPIs), making it easy to track progress and identify areas for improvement. Custom alerts and reports ensure that critical information is always at your fingertips.

**3. Continuous Learning and Improvement**

Cisco AppDynamics encourages continuous learning and improvement through its robust training resources and community support. Regular updates and enhancements keep the platform aligned with the latest technological advancements, ensuring that your APM strategy evolves alongside your business needs.

What is GKE-Native Monitoring?

GKE-Native Monitoring is a comprehensive monitoring solution provided by Google Cloud specifically designed for Kubernetes workloads running on GKE. It offers deep visibility into the performance and health of your clusters, allowing you to identify and address issues before they impact your applications proactively.

a) Automatic Cluster Monitoring: GKE-Native Monitoring automatically collects and visualizes key metrics from your GKE clusters, making monitoring your workloads’ overall health and resource utilization effortlessly.

b) Customizable Dashboards: With GKE-Native Monitoring, you can create personalized dashboards tailored to your specific monitoring needs. Visualize metrics that matter most to you and gain actionable insights at a glance.

c) Alerting and Notifications: Use GKE-Native Monitoring’s robust alerting capabilities to stay informed about critical events and anomalies in your GKE clusters. Configure alerts based on thresholds and receive notifications through various channels to ensure prompt response and issue resolution.

The Role of Network Baselining

Network baselining involves capturing and analyzing network traffic data to establish a benchmark or baseline for normal network behavior. This baseline represents the typical performance metrics of the network under regular conditions. It encompasses various parameters such as bandwidth utilization, latency, packet loss, and throughput. By monitoring these metrics over time, administrators can identify patterns, trends, and anomalies, enabling them to make informed decisions about network optimization and troubleshooting.

Understanding TCP Performance Parameters

TCP, or Transmission Control Protocol, is a vital protocol that governs reliable data transmission over networks. Behind its seemingly simple operation lies a complex web of performance parameters that can significantly impact network efficiency, latency, and throughput. In this blog post, we will dive deep into TCP performance parameters, understanding their importance and how they influence network performance.

TCP performance parameters determine various aspects of the TCP protocol’s behavior. These parameters include window size, congestion control algorithms, maximum segment size (MSS), retransmission timeout (RTO), and many more. Each parameter plays a crucial role in shaping TCP’s performance characteristics, such as reliability, congestion avoidance, and flow control.

The Impact of Window Size: Window size, also known as the receive window, represents the amount of data a receiving host can accept before requiring acknowledgment from the sender. A larger window size allows for more extensive data transfer without waiting for acknowledgments, thereby improving throughput. However, a vast window size can lead to network congestion and increased latency. Finding the optimal window size requires careful consideration and tuning.

Congestion Control Algorithms: Congestion control algorithms, such as TCP Reno, Cubic, and New Reno, regulate the flow of data in TCP connections to avoid network congestion. These algorithms dynamically adjust parameters like the congestion window and the slow-start threshold based on various congestion indicators. Understanding the different congestion control algorithms and selecting the appropriate one for specific network conditions is crucial for achieving optimal performance.

Maximum Segment Size (MSS): The Maximum Segment Size (MSS) refers to the most significant amount of data TCP can encapsulate within a single IP packet. It is determined by the underlying network’s Maximum Transmission Unit (MTU). A higher MSS can enhance throughput by reducing the overhead associated with packet headers, but it should not exceed the network’s MTU to avoid fragmentation and subsequent performance degradation.

Retransmission Timeout (RTO): Retransmission Timeout (RTO) is the duration at which TCP waits for an acknowledgment before retransmitting a packet. Setting an appropriate RTO value is crucial to balance reliability and responsiveness. A too-short RTO may result in unnecessary retransmissions and increased network load, while a too-long RTO can lead to higher latency and decreased throughput. Factors like network latency, jitter, and packet loss rate influence the optimal RTO configuration.

 

Baseline Engineering

Interactive Baseline Engineering Lab 02

Baseline Drift & Anomaly Investigation

Your baseline has already been established. Now investigate a live telemetry sample that looks different from normal. Identify the strongest evidence before choosing a likely cause.

Incident Scenario

Users report that a distributed application is responding slowly. Do not immediately blame the application. Compare current telemetry with the known baseline and determine where the strongest deviation appears.

Telemetry Comparison

Baseline vs Live
WAN Latency
22 ms
67 ms
Packet Loss
0.2%
0.3%
Throughput
640 Mbps
635 Mbps
App Response
180 ms
420 ms

Choose an Investigation Path

Evidence first
WAN / Network Path Check whether latency, jitter or loss has moved significantly from the established network baseline.
Application Dependency Check service-to-service latency, dependency response time and whether a downstream component is slowing the request path.
Capacity / Resource Pressure Check CPU, queue depth, throughput and other capacity indicators before concluding that the network is the problem.

Investigation Evidence

Awaiting selection
WAN Latency 22 ms → 67 ms
+204%
Packet Loss 0.2% → 0.3%
Small increase
Throughput 640 → 635 Mbps
Near baseline
Investigation console ready.
Baseline reference loaded.
Live telemetry sample loaded.

Investigation principle: A baseline identifies deviation; it does not automatically identify root cause. Correlate multiple signals and follow the dependency path before making a remediation decision.
ENGINEERING TAKEAWAY

Baseline comparison converts an unclear symptom into measurable evidence. When application latency rises, compare network, capacity and dependency signals together. The largest deviation is an investigation clue—not automatically the root cause.

Chaos Engineering

Chaos engineering is a methodology for experimenting with a software system to build confidence in its capability to withstand turbulent environments in production. It is an essential part of the DevOps philosophy, allowing teams to experiment with their system’s behavior in a safe and controlled manner.

This type of baseline engineering allows teams to identify weaknesses in their software architecture, such as potential bottlenecks or single points of failure, and take proactive measures to address them. By injecting faults into the system and measuring the effects, teams gain insights into system behavior that can be used to improve system resilience.

Finally, chaos Engineering teaches you to develop and execute controlled experiments that uncover hidden problems. For instance, you may need to inject system-shaking failures that disrupt system calls, networking, APIs, and Kubernetes-based microservices infrastructures.

Chaos engineering is “the discipline of experimenting on a system to build confidence in the system’s capability to withstand turbulent conditions in production.” In other words, it’s a software testing method that concentrates on finding evidence of problems before users experience them.

Network Baselining

Network baseline involves measuring the network’s performance at different times. This includes measuring throughput, latency, and other performance metrics and the network’s configuration. It is important to note that performance metrics can vary greatly depending on the type of network being used. This is why it is essential to establish a baseline for the network to be used as a reference point for comparison.

Network baselining is integral to network management. It allows organizations to identify and address potential issues before they become more serious. Organizations can be alerted to potential problems by analyzing the network’s performance. This can help organizations avoid costly downtime and ensure their networks run at peak performance.

network baselining
Diagram: Network Baselining. Source is DNSstuff

**The Importance of Network Baselining**

Network baselining provides several benefits for network administrators and organizations:

1. Performance Optimization: Baselining helps identify bottlenecks, inefficiencies, and abnormal behavior within the network infrastructure. By understanding the baseline, administrators can optimize network resources, improve performance, and ensure a smoother user experience.

2. Security Enhancement: Baselining also plays a crucial role in detecting and mitigating security threats. Administrators can identify unusual or malicious activities by comparing current network behavior against the established baseline, such as abnormal traffic patterns or unauthorized access attempts.

3. Capacity Planning: Understanding network baselines enables administrators to forecast future capacity requirements accurately. By analyzing historical data, they can determine when and where network upgrades or expansions may be necessary, ensuring consistent performance as the network grows.

**Establishing a Network Baseline**

To establish an accurate network baseline, administrators follow a systematic approach:

1. Data Collection: Network traffic data is collected using specialized monitoring tools like network analyzers or packet sniffers. These tools capture and analyze network packets, providing detailed insights into performance metrics.

2. Duration: Baseline data should ideally be collected over an extended period, typically from a few days to a few weeks. This ensures the baseline accounts for variations due to different network usage patterns.

3. Normalizing Factors: Administrators consider various factors impacting network performance, such as peak usage hours, seasonal variations, and specific application requirements. Normalizing the data can establish a more accurate baseline that reflects typical network behavior.

4. Analysis and Documentation: Once the baseline data is collected, administrators analyze the metrics to identify patterns and trends. This analysis helps establish thresholds for acceptable performance and highlights any deviations that may require attention. Documentation of the baseline and related analysis is crucial for future reference and comparison.

Network Baselining: A Lot Can Go Wrong

Infrastructure is becoming increasingly complex, and let’s face it, a lot can go wrong. It’s imperative to have a global view of all the infrastructure components and a good understanding of the application’s performance and health. In a large-scale container-based application design, there are many moving pieces and parts, and it is hard to validate the health of each piece manually.  

Therefore, monitoring and troubleshooting are much more complex, especially as everything is interconnected, making it difficult for a single person in one team to understand what is happening entirely. Nothing is static anymore; things are moving around all the time. This is why it is even more important to focus on the patterns and to be able to see the path of the issue efficiently.

Some modern applications could simultaneously be in multiple clouds and different location types, resulting in numerous data points to consider. If any of these segments are slightly overloaded, the sum of each overloaded segment results in poor performance on the application level. 

What does this mean to latency?

Distributed computing has many components and services, with far-apart components. This contrasts with a monolith, with all parts in one location. Because modern applications are distributed, latency can add up. So, we have both network latency and application latency. The network latency is several orders of magnitude more significant.

As a result, you need to minimize the number of Round-Trip Times and reduce any unneeded communication to an absolute minimum. When communication is required across the network, it’s better to gather as much data together as possible to get bigger packets that are more efficient to transfer. Also, consider using different types of buffers, both small and large, which will have varying effects on the dropped packet test.

With the monolith, the application is simply running in a single process, and it is relatively easy to debug. Many traditional tooling and code instrumentation technologies have been built, assuming you have the idea of a single process. The core challenge is trying to debug microservices applications. So much of the tooling we have today has been built for traditional monolithic applications. So, there are new monitoring tools for these new applications, but there is a steep learning curve and a high barrier to entry.

A new approach: Network baselining and Baseline engineering

For this, you need to understand practices like Chaos Engineering, along with service level objectives (SLOs), and how they can improve the reliability of the overall system. Chaos Engineering is a baseline engineering practice that allows tests to be performed in a controlled way. Essentially, we intentionally break things to learn how to build more resilient systems.

So, we are injecting faults in a controlled way to make the overall application more resilient by injecting various issues and faults. Implementing practices like Chaos Engineering will help you understand and manage unexpected failures and performance degradation. The purpose of Chaos Engineering is to build more robust and resilient systems.

**A final note on baselines: Don’t forget them**

Creating a good baseline is a critical factor. You need to understand how things work under normal circumstances. A baseline is a fixed point of reference used for comparison purposes. You usually need to know how long it takes to start the application to the actual login and how long it takes to do the essential services before there are any issues or heavy load. Baselines are critical to monitoring.

It’s like security; if you can’t see what, you can’t protect. The same assumptions apply here. Go for a good baseline and if you can have this fully automated. Tests need to be carried out against the baseline on an ongoing basis. You need to test constantly to see how long it takes users to use your services. Without baseline data, estimating any changes or demonstrating progress is difficult.

Network baselining is a critical practice for maintaining optimal network performance and security. By establishing a baseline, administrators can proactively monitor, analyze, and optimize their networks. This approach enables them to promptly identify and address performance issues, enhance security measures, and plan for future capacity requirements. Organizations can ensure a reliable and efficient network infrastructure that supports their business objectives by investing time and effort in network baselining.

Interactive Baseline Engineering Lab 03

Chaos Engineering & Baseline Validation

A baseline is only useful if it survives controlled testing. Inject a fault, observe how telemetry changes, recover the system and decide whether the original baseline still represents normal behaviour.

Scenario

The production baseline says the service normally operates with low latency, minimal packet loss and predictable application response time. You will deliberately introduce controlled failure conditions and observe the resulting telemetry. This is a teaching simulation of chaos engineering—not a production fault-injection system.

Controlled Failure Injection

Healthy
Client Requests
Load Balancer Traffic
Service Application
Database Dependency
Network Latency Introduce additional path delay.
Packet Loss Simulate degraded network delivery.
Traffic Increase Push the service beyond its normal workload.
Instance Failure Remove an application instance from service.
Database Delay Increase downstream database response time.

Live Telemetry

Compared with baseline
WAN Latency
22 ms
22 ms
Packet Loss
0.2%
0.2%
Throughput
640 Mbps
640 Mbps
CPU
48%
48%
App Response
180 ms
180 ms
Baseline
Fault Injected
Observe
Recover

Experiment Console

Ready
Experiment environment initialised.
Known-good baseline loaded.
Waiting for controlled fault injection.

ENGINEERING TAKEAWAY

Chaos engineering deliberately introduces controlled failure to test system behaviour. Baseline engineering gives you the reference point needed to determine what changed, whether recovery worked and whether the original definition of “normal” still reflects the system.

Docker network security

Docker Security Options

Docker Security Options

Container security is no longer limited to protecting a Docker daemon or scanning an image for known vulnerabilities. Modern container environments span the software supply chain, container runtime, host operating system, network, secrets, registries, CI/CD pipelines, and often an orchestration platform such as Kubernetes.

Docker provides a powerful packaging and execution model, but containers are not equivalent to virtual machines. Containers share the host kernel, which means a compromised or misconfigured container can potentially expose the underlying host or other workloads. Effective container security therefore requires controls at multiple layers.

Understanding Container Security Risks

Common risks include vulnerable base images, malicious or tampered dependencies, exposed container services, excessive Linux capabilities, insecure filesystem permissions, leaked credentials, privileged containers, container breakout vulnerabilities, and compromised registries or build pipelines.

Security should therefore begin before a container reaches production. An image that is already compromised or contains a known critical vulnerability cannot be made secure simply by applying runtime controls later.

Image and Supply Chain Security

Container images should be built from trusted and maintained base images and kept as small as practical. Reducing unnecessary packages decreases the attack surface and limits the number of components requiring security updates.

Image scanning can identify known vulnerabilities in operating-system packages and application dependencies. However, vulnerability scanning should be combined with software supply-chain controls such as trusted registries, dependency management, image provenance, cryptographic signing or attestations, and controlled build pipelines.

Image tags such as latest should not be treated as a strong integrity mechanism because they can change over time. Production workflows can instead use immutable image references or digests so that the exact image being deployed is known.

Least Privilege at Runtime

A container should receive only the privileges required for its workload. Running containers as root when it is unnecessary increases the potential impact of a compromise.

Security controls can include non-root users, restricted Linux capabilities, read-only filesystems where practical, controlled device access, restricted privilege escalation, seccomp profiles, and mandatory access controls such as AppArmor or SELinux where supported.

Privileged containers deserve particular scrutiny because they can significantly weaken the isolation boundary between the workload and the host.

Secrets and Credentials

Secrets should never be embedded directly into container images or source repositories. Database passwords, API tokens, certificates, and other credentials should be managed through appropriate secret-management mechanisms and injected into workloads only when required.

Access to secrets should follow least-privilege principles, with appropriate rotation, auditing, and separation between development, testing, and production environments.

Container Network Security

Container networking can create additional attack paths if workloads are allowed unrestricted communication. Network segmentation should therefore be designed around application requirements rather than simply placing every container on the same network.

Docker networks, host firewalls, cloud security controls, and, in orchestrated environments, network policies can be used to restrict which workloads can communicate with one another and which services are exposed externally.

Encryption should also be considered whenever sensitive information crosses networks or leaves the host.

Runtime and Host Security

Container security depends heavily on the underlying host. Docker Engine, containerd, the Linux kernel, supporting packages, and host configuration must all be maintained and monitored.

The container runtime provides isolation mechanisms such as Linux namespaces and control groups, while additional security mechanisms can restrict system calls, capabilities, filesystem access, and interaction with the host.

A vulnerable host can undermine otherwise well-designed container controls, making host patching and configuration management an essential part of the overall security model.

Monitoring and Detection

Security does not end when a container starts successfully. Runtime activity should be monitored for unexpected processes, unusual network connections, privilege changes, filesystem modifications, resource anomalies, and other indicators of compromise.

Container logs, host telemetry, runtime events, vulnerability reports, and security alerts can be combined with broader observability and security monitoring platforms to provide operational context.

Continuous Container Security

Container security is a lifecycle rather than a single configuration task:

Build securely → Scan → Verify provenance → Store securely → Deploy with least privilege → Restrict network access → Monitor → Patch and rebuild

This approach recognises that vulnerabilities will continue to emerge in operating-system packages, application dependencies, runtimes, and the underlying infrastructure.

The objective is not to make a container permanently “secure.” Instead, organisations should build a repeatable security process that reduces attack surface, limits privileges, protects the supply chain, detects abnormal behaviour, and enables vulnerable workloads to be rebuilt and replaced quickly.

A strong Docker security strategy therefore treats the container as one component of a larger system. Image integrity, runtime isolation, host security, identity, secrets, networking, monitoring, and continuous vulnerability management must work together to protect the workload throughout its lifecycle.

Highlights: Docker Security Options

Interactive Security Lab

Docker Security Posture Builder

Start with a deliberately weak container configuration and progressively apply defense-in-depth controls. Your goal is to reduce privilege, constrain the runtime, protect the image and limit the blast radius.

Security posture
15 / 100
CRITICAL
Container is broadly exposed

Identity & Privilege

Reduce what the container can do if compromised.
Run as non-root Remove unnecessary root privileges inside the container.
High
Rootless mode Run Docker without requiring root-level daemon privileges.
High
Drop unnecessary capabilities Minimize Linux capabilities such as CAP_SYS_ADMIN.
High
No New Privileges Prevent processes from gaining additional privileges.
Med

Runtime Isolation

Constrain system calls, filesystems and resources.
Read-only filesystem Prevent unnecessary writes to the container filesystem.
Med
Seccomp profile Restrict access to unnecessary or dangerous system calls.
High
AppArmor / SELinux policy Apply mandatory access controls to constrain processes.
Med
CPU / memory / PID limits Use resource controls to reduce exhaustion and abuse.
Med

Image & Supply Chain

Reduce the chance that the workload starts from an unsafe image.
Scan container images Identify known vulnerabilities before deployment.
High
Verify image signing / provenance Establish trust in where images came from.
Med
Use a minimal base image Reduce unnecessary packages and attack surface.
Med

Network & Data Protection

Limit connectivity and protect sensitive workload data.
Network isolation Allow only the connectivity the workload actually requires.
High
Use secrets management Keep credentials and sensitive values out of images.
Med
Health checks Provide a defined signal for workload health and recovery.
Low

Hardened configuration

0 controls enabled
Critical current risk level
0% posture coverage
# Container starts with weak defaults
docker run my-app:latest
# root + broad privileges + writable filesystem
Engineering takeaway: Container security is defense in depth. Non-root execution, reduced capabilities, syscall filtering, mandatory access controls, filesystem restrictions, resource limits, trusted images and network isolation address different parts of the attack surface. No single control makes a container secure by itself.

Docker Security Options

Docker Security Options 

Core Docker Security Options

  • User namespaces

  • Seccomp profiles

  • AppArmor / SELinux

  • Linux capabilities

  • Read‑only filesystem

  • Rootless mode

  • cgroups resource limits

  • No New Privileges

  • Image signing

  • Network isolation

  • Secrets management

  • Health checks

 

1. Linux Security Features

User Namespaces

Maps container root to an unprivileged host user, reducing breakout risk.

Seccomp

Restricts dangerous system calls. Docker’s default profile blocks many high‑risk syscalls.

AppArmor / SELinux

Mandatory access control that limits filesystem access, kernel interactions, and process behavior.

Capabilities

Drop unnecessary Linux capabilities (e.g., NET_ADMIN, SYS_ADMIN) to reduce privilege.

 

2. Container Runtime Hardening

Rootless Mode

Runs Docker without host root privileges.

Read‑Only Filesystem

Prevents file tampering inside containers.

No New Privileges

Blocks privilege escalation even if binaries have setuid bits.

cgroups

Limits CPU, memory, I/O, and PIDs to prevent resource exhaustion.

 

3. Image Security

Image Scanning

Scan images for vulnerabilities using tools like Trivy, Clair, or Anchore.

Image Signing

Verify trusted images using Notary or Sigstore.

Minimal Base Images

Use Alpine, Distroless, or Scratch to reduce attack surface.

 

4. Network Security

Isolated Networks

Use custom bridge or overlay networks with strict rules.

Firewall Rules

Restrict container ingress and egress.

Service Mesh

Add mTLS, identity‑aware routing, and policy enforcement.

 

5. Runtime Monitoring

Use tools like Falco, Sysdig, Prometheus, or OpenTelemetry to monitor:

  • syscalls

  • network flows

  • container lifecycle

  • anomalies

Outcome

Docker security options provide:

  • strong isolation

  • reduced breakout risk

  • hardened runtime

  • secure images

  • controlled network behavior

  • predictable resource usage

  • compliance with enterprise security standards

Container Security

The fact that containers share the kernel of the Linux server boosts their performance and makes them lightweight. Because of this, Linux containers pose the most significant security risk. Namespaces are not everywhere in the kernel, which is the main reason for this concern.

Because cgroups and standard namespaces provide some necessary isolation from the host’s core resources, containerized applications are more secure than noncontainerized applications. However, containers should not be used as a replacement for good security practices. It would be best if you run all your containers as you would run an application on a production system. The same should apply if your application runs as a nonprivileged user on a server.

Docker Attack Surface

So you are currently in the Virtual Machine world and considering transitioning to a containerized environment. You want to smoothen your application pipeline and gain the benefits of a Docker containerized environment. But you have heard from many that the containers are insecure and are concerned about Docker network security. There is a Docker attack surface to be concerned about.

Example: Containers run as root

For example, containers run by root by default and have many capabilities that scare you. Yes, we have a lot of benefits to the containerized environment, and containers are the only way to do it for some application stacks. However, we have a new attack surface with some benefits of deploying containers and forcing you to examine Docker security options. The following post will discuss security issues, a container security video to help you get started, and an example of Docker escape techniques.

Understanding SELinux

SELinux, which stands for Security-Enhanced Linux, is a security framework built into the Linux kernel. It provides a powerful set of security policies and access controls to enforce fine-grained restrictions on processes and resources. By leveraging SELinux, administrators can define and enforce access rules, reducing the attack surface of Docker containers.

It is important to understand how SELinux and Docker can be integrated to utilize SELinux in a Docker environment. Docker provides SELinux support through SELinux labels applied to Docker objects such as containers and volumes. These labels enforce SELinux policies and restrict container actions based on defined policies.

Advantages: SELinux

SELinux offers several benefits for Docker security. Firstly, it provides mandatory access controls, enabling administrators to define precisely what actions a container can perform. This prevents malicious containers from accessing sensitive resources or executing unauthorized commands. Secondly, SELinux helps mitigate container breakout attacks by isolating containers and limiting their interactions with the host system. Lastly, SELinux can help detect and prevent privilege escalation attempts within Docker containers.

**New Attacks and New Components**

Containers are secure by themselves, and the kernel is pretty much battle-tested. A container escape is hard to orchestrate unless misconfiguration could result in excessive privileges. So, even though the bad actors’ intent may stay the same, we must mitigate a range of new attacks and protect new components.

To combat these, you need to be aware of the most common Docker network security options and follow the recommended practices for Docker container security. A platform approach is also recommended, and OpenShift is a robust platform for securing and operating your containerized environment.

Understanding Docker Bench Security

Docker Bench Security, developed by the Center for Internet Security (CIS), is a script that automates the process of running security checks against Docker installations. It follows industry-standard best practices and provides a comprehensive report on potential vulnerabilities and misconfigurations.

Section 1: Installation and Configuration

To begin using Docker Bench Security, you first need to install it on your host system. The installation process is straightforward and well-documented. Once installed, you can configure the script to perform specific checks based on your needs and environment.

Section 2: Running Docker Bench

Running Docker Bench Security is as simple as executing a single command. The script will systematically analyze your Docker setup, checking for various security aspects such as host configuration, Docker daemon configuration, container runtime, networking, and more. It will generate a detailed report highlighting any security issues found.

The generated report from Docker Bench Security provides valuable insights into your Docker environment’s security posture. It categorizes the findings into different levels of severity, helping you prioritize and address the most critical vulnerabilities first. By understanding the report and taking necessary actions, you can significantly enhance the security of your Docker deployments.

 

Interactive Incident Investigation

Docker Attack Surface Investigation Lab

A container has been compromised. Investigate the configuration evidence and identify the conditions that allowed the attacker to increase impact beyond the container itself.

Scenario: An application container was deployed from an untrusted image. The workload is running with excessive privileges and has broad network connectivity. Your task is to identify the highest-risk findings and reconstruct the potential attack path.

1. Examine the evidence

Select the findings that materially increase the container's attack surface.
IMG
Untrusted / vulnerable image Image provenance is unknown and packages contain known weaknesses.
SELECT
CAP
CAP_SYS_ADMIN present Powerful Linux capability increases the consequences of compromise.
SELECT
FS
Writable filesystem Compromised processes can modify files inside the container.
SELECT
NET
Broad east-west connectivity The workload can communicate with services it does not require.
SELECT
LOG
Application logging enabled Logging provides useful visibility but is not itself an attack vector.
SELECT

2. Reconstruct the attack path

The selected findings reveal how a container compromise can increase impact.
Untrusted image
↓
Vulnerable application / process
↓
Excessive runtime privileges
↓
Shared host kernel boundary
↓
Potential host / lateral impact
Investigation status
Evidence incomplete
Select the relevant security findings, then run the investigation.
Engineering takeaway: Containers provide process and resource isolation, but they share the host kernel. Excessive privileges, dangerous capabilities, vulnerable images and unnecessary network access can increase the impact of a compromise. The objective is not to assume a container is invulnerable; it is to reduce both the probability and blast radius of compromise.

Docker Security Options

Docker Security

To use Docker safely in production and development, you must know potential security issues and the primary tools and techniques for securing container-based systems. Your system’s defenses should also consist of multiple layers.

For example, your containers will most likely run in VMs so that if a container breakout occurs, another level of defense can prevent the attacker from getting to the host or other containers. Monitoring systems should be in place to alert admins in the case of unusual behavior. Finally, firewalls should restrict network access to containers, limiting the external attack surface.

Container Isolation:

One of Docker’s key security features is container isolation, which ensures that each container runs in its own isolated environment. By utilizing Linux kernel features such as namespaces and cgroups, Docker effectively isolates containers from each other and the host system, mitigating the risk of unauthorized access or interference between containers.

Image Vulnerability Scanning:

It is crucial to scan Docker images for vulnerabilities regularly to ensure their security. Docker Security Scanning is an automated service that helps identify known security issues in your containers’ base images and dependencies. By leveraging this feature, you can proactively address vulnerabilities and apply necessary patches, reducing the risk of potential exploits.

Docker Content Trust:

Docker Content Trust is a security feature that allows you to verify the authenticity and integrity of images you pull from Docker registries. By enabling this feature, Docker ensures that only signed and verified images are used, preventing the execution of untrusted or tampered images. This provides an additional layer of protection against malicious or compromised containers.

Role-Based Access Control (RBAC):

Controlling access to Docker resources is critical to maintaining a secure environment. Docker Enterprise Edition (EE) offers Role-Based Access Control (RBAC), which allows you to define granular access controls for users and teams. By assigning appropriate roles and permissions, you can restrict access to sensitive operations and ensure that only authorized individuals can manage Docker resources.

Network Segmentation:

Docker provides various networking options to facilitate communication between containers and the outside world. Implementing network segmentation techniques, such as bridge or overlay networks, helps isolate containers and restrict unnecessary network access. By carefully configuring the network settings, you can minimize the attack surface and protect your containers from potential network-based threats.

Container Runtime Security:

In addition to securing the container environment, it is equally important to focus on the security of the container runtime. Docker supports different container runtimes, such as Docker Engine and containerd. Regularly updating these runtimes to the latest stable versions ensures that you benefit from the latest security patches and bug fixes, reducing the risk of potential vulnerabilities.

**Docker Attack Surface**

Often, the tools and appliances in place are entirely blind to containers. The tools look at a running process and think, if the process is secure, then I’m safe. One of my clients ran a container with the DockerFile and pulled an insecure image. The onsite tools did not know what an image was and could not scan it.

As a result, we had malware right in the network’s core, a little bit too close to the database server for my liking.  Yes, we call containers a fancy process, and I’m to blame here, too, but we need to consider what is around the container to secure it fully. For a container to function, it needs the support of the infrastructure around it, such as the CI/CD pipeline and supply chain.

To improve your security posture, you must consider all the infrastructures. If you are looking for quick security tips on Docker network security, this course I created for Pluralsight may help you with Docker security options.

Ineffective Traditional Tools

Containers are not like traditional workloads. With a single command, we can run an entire application with all its dependencies. Legacy security tools and processes often assume largely static operations and must be adjusted to adapt to the rate of change in containerized environments. With non-cloud-native data centers, Layer 4 is coupled with the network topology at fixed network points and lacks the flexibility to support containerized applications.

There is often only inter-zone filtering and east-to-west traffic may go unchecked. A container changes the perimeter and moves right to the workload. Just look at a microservices architecture. It has many entry points compared to monolithic applications.

Docker container networking

When considering container networking, we are a world apart from the monolithic. Containers are short-lived and constantly spun down, and assets such as servers, IP addresses, firewalls, drives, and overlay networks are recycled to optimize utilization and enhance agility. Traditional perimeters designed with I.P. address-based security controls lag in a containerized environment.

Rapidly changing container infrastructure rules and signature-based controls can’t keep up with a containerized environment. Securing hyper-dynamic container infrastructure using traditional networks ​​and endpoint controls won’t work. For this reason, you should adopt purpose-built tools and techniques for a containerized environment.

**The Need for Observability**

Not only do you need to implement good Docker security options, but you also need to be concerned about the recent observability tools. So, we need proper observability of the state of security and the practices used in the containerization environment, and we need to automate this as much as possible—not just the development but also the security testing, container scanning, and monitoring.

You are only as secure as the containers you have running. You need to be observable in systems and applications and proactive in these findings. It is not something you can buy; it is a cultural change. You want to know how the application works with the server, how the network is with the application, and what data transfer looks like in transfer and a stable state.  

What level of observation do you need so you know that everything is performing as it should? There are several challenges to securing a containerized environment. Containerized technologies are dynamic and complex and require a new approach that can handle the agility and scale of today’s landscape. There are initial security concerns that you must understand before you get started with container security. This will help you explore a better starting strategy.

Docker attack surface: Container attack vectors 

We must consider a different threat model and understand how security principles such as least privilege and in-depth defense apply to Docker security options. With Docker containers, we have a completely different way of running applications and, as a result, a different set of risks to deal with.

Instructions are built into Dockerfiles, which run applications differently from a normal workload. With the correct rights, a bad actor could put anything in the Dockerfile without the necessary guard rails that understand containers; there will be a threat.

Therefore, we must examine new network and security models, as old tools and methods won’t meet these demands.  A new network and security model requires you to mitigate against a new attack vector. Bad actors’ intent stays the same. They are not going away anytime soon. But they now have a different and potentially easier attack surface if misconfigured.

I would consider the container attack surface pretty significant; bad actors will have many default tools if not locked down. For example, we have image vulnerabilities, access control exploits, container escapes, privilege escalation, application code exploits, attacks on the docker host, and all the docker components.

Docker security options: A final security note

Containers by themselves are secure, and the kernel is pretty much battle-tested. You will not often encounter kernel compromises, but they happen occasionally. A container escape is brutal to orchestrate unless misconfiguration could result in excessive privileges. From a security standpoint, it would be best to stay clear of setting container capabilities that provide excessive privileges.

Minimise container capabilities: Reduce the attack surface.

If you minimize the container’s capabilities, you are stripping down its functionality to a bare minimum—we mentioned this in the container security video. Therefore, the attack surface is limited, and the attack vector available to the attacker is minimized. 

You also want to keep an eye on CAP_SYS_ADMIN. This flag grants access to an extensive range of privileged activities. Containers run many other capacities by default that can cause havoc.

As Docker continues to gain popularity, understanding and implementing proper security measures is essential to safeguarding your containers and infrastructure. By leveraging the security options discussed in this blog post, you can mitigate risks, protect against potential threats, and ensure the integrity and confidentiality of your applications. Stay vigilant, stay secure, and embrace the power of Docker while keeping your containers safe.

Interactive Containment Lab

Container Network Segmentation Lab

Design least-privilege connectivity between application tiers, then test what happens when a frontend container is compromised and attempts lateral movement toward the database.

Target application topology

Internet External traffic
→
Firewall Policy boundary
→
Frontend Web tier
→
API Application tier
→
Database Data tier

Required connectivity

A least-privilege design allows only the communication required by the application architecture.
Internet → Frontend Public web access
ALLOW
Frontend → API Application requests
ALLOW
API → Database Required data access
ALLOW
Frontend → Database No direct database access required
DENY
Database → Internet Database should not initiate external access
DENY
API → Internet Restrict unless explicitly required
DENY

Test the containment boundary

A compromised frontend attempts to bypass the API and reach the database directly.
⚠ Compromised Frontend

Attacker attempts: Frontend → Database

Segmentation result
Ready to test
Run the connectivity test to determine whether lateral movement is contained.
Engineering takeaway: Network isolation should follow application dependencies. The frontend does not need direct database connectivity simply because both workloads exist on the same infrastructure. Restricting unnecessary paths reduces lateral movement and limits blast radius after compromise. Docker networking mechanisms and Kubernetes NetworkPolicy are different technologies; the security principle here is least-privilege network connectivity.