Kubernetes PetSets
In Kubernetes, managing applications that maintain persistent state requires different considerations from running stateless workloads. Web front ends and many API services can often be replaced or rescheduled without preserving the identity of an individual Pod. Databases, distributed queues, clustered storage systems, and other stateful workloads may require persistent storage, stable network identity, predictable startup behaviour, and controlled updates.
The Kubernetes resource designed for this purpose is the StatefulSet. StatefulSet was originally introduced as PetSet and was renamed as Kubernetes evolved. The current StatefulSet API provides a declarative mechanism for managing Pods that require persistent identity and storage while still benefiting from Kubernetes scheduling, service discovery, health management, and automation.
A StatefulSet gives each Pod a predictable identity based on its ordinal position. For example, a StatefulSet with three replicas might create Pods named database-0, database-1, and database-2. These identities remain associated with the corresponding Pods even when individual instances are rescheduled. This predictable identity is particularly useful for distributed systems where members need to discover and distinguish one another.
Persistent storage is another fundamental StatefulSet capability. StatefulSets can use volumeClaimTemplates to create persistent volume claims associated with individual Pod identities. When a Pod is recreated, Kubernetes can reattach the appropriate persistent storage rather than treating the replacement as an entirely new stateless instance. The actual durability and availability characteristics depend on the underlying StorageClass and storage platform.
StatefulSets also provide controlled Pod management. Depending on the configured policy, Pods can be created, updated, or terminated in an ordered manner. This can be important for clustered applications that need members to start in a particular sequence or require controlled membership changes. However, ordering should not be confused with application-level consistency; the application itself remains responsible for replication, quorum, transactions, and data integrity.
Stable network identity is commonly provided through a Kubernetes Service, often a headless Service, combined with the StatefulSet’s predictable Pod identities. This allows distributed applications to discover individual members rather than communicating only with an interchangeable load-balanced endpoint. DNS-based service discovery can then provide stable names even as Pods are rescheduled.
StatefulSets support rolling updates and recovery mechanisms, but they do not automatically make a stateful application highly available. Kubernetes can restart or reschedule a failed Pod, but database replication, leader election, quorum management, backup, recovery, and consistency are responsibilities of the application or an appropriate Kubernetes operator. For complex databases and distributed systems, operators can provide application-specific automation that goes beyond the generic StatefulSet abstraction.
Storage design is therefore a critical part of any StatefulSet deployment. Engineers should consider the StorageClass, access modes, performance characteristics, failure domains, backup strategy, snapshot capabilities, encryption, and recovery procedures. Persistent volumes provide storage persistence, but persistence is not the same as backup. A resilient stateful platform needs a tested recovery strategy in addition to persistent storage.
Observability is equally important. Stateful workloads should be monitored for Pod health, storage capacity, I/O performance, replication status, application errors, restart frequency, resource consumption, and cluster-specific conditions. Kubernetes metrics and events can provide infrastructure-level information, while application telemetry is often required to understand whether the database or distributed system itself is operating correctly.
StatefulSets are therefore best understood as a Kubernetes mechanism for providing stable identity, persistent storage relationships, and controlled lifecycle management. They are not a generic solution that turns any stateful application into a distributed, fault-tolerant service.
The modern Kubernetes approach is to combine StatefulSets with persistent storage, Services, appropriate scheduling and disruption policies, observability, backup and recovery processes, and—where appropriate—application-specific operators. Used correctly, this provides a strong foundation for running stateful workloads while preserving the declarative and automated operating model that makes Kubernetes powerful.
Matt Conran
Pets vs Cattle Lab
See why Kubernetes treats ordinary stateless replicas differently from stateful workloads that need persistent identity, storage and predictable lifecycle behaviour.
Highlights: Kubernetes PetSets
Stateful Applications with Kubernetes PetSets
a) PetSets, introduced in Kubernetes version 1.3, provides a higher-level abstraction for managing stateful applications. Unlike traditional Kubernetes deployments, which focus on stateless workloads, PetSets offer features like stable network identities, ordered deployment, and automated scaling while considering the application’s stateful nature.
b) One key challenge in scaling stateful applications is ensuring stable network identities. PetSets addresses this by providing stable hostnames and domain names for each replica of the stateful application. This allows clients to consistently connect to the same replica, even when scaling or restarting instances.
c) PetSets enable ordered deployment and scaling of stateful applications. This is crucial when dealing with applications that rely on specific ordering or coordination between instances. With PetSets, you can define a startup order for your replicas, ensuring that each replica is fully operational before the next one starts.
Benefits of PetSets
PetSets offer several advantages over traditional Kubernetes deployments:
1. Stable Network Identity: PetSets assigns each pod a stable hostname and DNS identity, enabling seamless communication between pods and external services. This stability is essential for applications relying on peer-to-peer communication or distributed databases.
2. Ordered Deployment: PetSets ensure the ordered deployment and scaling of pods. This is particularly useful for applications where the order of creation or scaling matters, such as databases or distributed systems.
3. Persistent Volumes: PetSets facilitate the association of persistent volumes with pods, ensuring data durability and allowing applications to retain their state even if the pods are rescheduled.
4. Rolling Updates: PetSets supports rolling updates, enabling you to update your stateful applications without downtime. This process ensures that each pod is gracefully terminated and replaced by a new one with the updated configuration, minimizing service interruptions.
Understanding Kubernetes PetSet
A – ) Kubernetes PetSet, also known as StatefulSet in more recent versions, is designed to manage stateful applications that require stable network identities and persistent storage. Unlike traditional Kubernetes Deployments, PetSet ensures that each pod in a set has a unique and stable hostname, allowing applications to maintain their identity and communicate effectively with other components in the cluster.
B – ) PetSet offers several key features that make it a valuable tool for managing stateful applications. One of its primary benefits is the ability to automatically provision and manage persistent volumes for each pod in the set. This ensures data durability and allows applications to seamlessly recover from pod failures without losing critical data. Additionally, PetSet provides ordered pod creation and termination, allowing for smooth scaling and rolling updates while preserving the application’s state.
C – ) Kubernetes PetSet finds application in various scenarios where stateful workloads are involved. One common use case is running databases, such as MySQL or PostgreSQL, in a distributed fashion. PetSet ensures that each replica of the database has a stable and unique identity, enabling seamless replication and failover. Other use cases include running distributed file systems, message queues, and other stateful applications that require stable network identities and persistent storage.
D – ) To effectively utilize Kubernetes PetSet, it’s essential to follow some best practices. Firstly, carefully design your application to ensure it can handle pod failures and rescheduling. Leveraging persistent volumes and configuring appropriate storage classes is crucial for data durability and availability. Additionally, monitoring the health and performance of your PetSet pods and implementing proper scaling strategies will help optimize the overall performance of your stateful application.
PetSets Use Cases: –
PetSets are ideal for various stateful applications, including databases, distributed systems, and legacy applications that require stable network identities and persistent storage. Some everyday use cases for PetSets include:
1. Running a replicated database cluster, such as MySQL or PostgreSQL, where each pod corresponds to a separate database node.
2. Managing distributed messaging systems, like Apache Kafka or RabbitMQ, where each pod represents a separate broker or message queue.
3. Deploying legacy applications that rely on stable network identities and persistent storage, such as content management systems or enterprise resource planning software.
**Example: Detera and stateful applications**
In Kubernetes, persistent volumes are critical as customers migrate from stateless workloads to stateful applications. Pet Sets have significantly improved Kubernetes’ support for stateful applications like MySQL, Kafka, Cassandra, and Couchbase. It was possible to automate the scaling of the “Pets” (applications that need persistent placement and consistent handling) using sequencing provisioning and startup procedures.
The Role of FlexVolume:
Datera integrates seamlessly with Kubernetes using FlexVolume, an elastic block storage system for cloud deployments. Based on the first principles of containers, Datera decouples the provisioning of application resources from the underlying physical infrastructure. Clean contracts (i.e., not dependent on physical infrastructure), declarative formats, and declarative formats can eventually make stateful applications portable.
YAML Configurations:
With Datera, Kubernetes allows the underlying application infrastructure to be defined through YAML configurations, which are passed to the storage infrastructure. Datara AppTemplates can automate the scaling of stateful applications in a Kubernetes environment.
**The Role of Kubernetes**
Kubernetes Pets and PetSets are core components of Kubernetes operations. Firstly, Kubernetes is a container orchestration platform that runs and manages containers. It changes the focus of container deployment to an application level, not the machine. The shift of focus point enables an abstraction level and the removal of dependencies between the application and its physical deployment.
**The Role of Decoupling**
This act of decoupling services from the details of low-level physical deployment enables better service management. For anything to scale, you need to provide some abstraction. For container networking, we have seen this in the overlay world with underlay and abstraction enabling networks to support millions of tenants.
**Kubernetes Networking 101**
Kubernetes Networking 101 allows the deployment of applications to a “sea of abstracted computes,” enabling a self-healing orchestrated infrastructure. While this scaling and deployment have been helpful for stateless services, they fall short in the stateful world with the base Kubernetes distribution. Most of this has been solved today with Red Hat products, including OpenShift networking, which has several new network and security constructs that aid with stateful application support.
StatefulSet Identity & Scaling Lab
Scale a stateful workload and watch Kubernetes preserve ordinal identity, persistent-volume relationships and predictable lifecycle behaviour.
Kubernetes PetSets
Discussing StatefulSets
Kubernetes started providing a resource to manage stateful workloads with the alpha release of PetSets in the 1.3 release. This capability has matured and is now known as StatefulSets. A StatefulSet has some similarities to a ReplicaSet in that it is responsible for managing the lifecycle of a set of Pods, but how it goes about this management has some noteworthy differences. PetSet might seem like an odd name for a Kubernetes resource, and it has since been replaced.
Still, it provides fascinating insights into the Kubernetes community’s thought process for supporting stateful workloads. The fundamental idea is that there are two ways of handling servers: to treat them as pets that require care, feeding, and nurturing or to treat them as cattle to which you don’t develop an attachment or provide much individual attention. If you’re logging into a server regularly to perform maintenance activities, you treat it as a pet.
**Contrasting Application Models**
As applications serve a growing user base around the globe, single cluster and data center solutions are no longer satisfying. Clustered applications and federations enable workloads to spread across multiple locations and container clusters for improved efficiency and scale. You will find contrasting deployment models when you examine the application types and scaling modes (for example, scale-up and scale-out clustering ).
There are substantial differences between deploying and running single applications to applications that operate within a cluster. Different application architectures require different deployment solutions and network identities. For example, a database node requires persistent volumes, or a node within a cluster uses specific elections where identity is essential.
**Kubernetes Pets **
Kubernetes has recently ramped up by introducing a new Kubernetes object called PetSets. PetSets is geared towards improving stateful support and is currently an alpha resource in Kubernetes release 1.3. It holds a group of Kubernetes Pets, aka stateful applications that require stable hostnames, persistent disks, a structured lifecycle process, and group identity.
PetSets are used for non-homogenous instances where each POD has a stable distinguishable identity. A different identity is viewed in terms of stable network and storage.
A ) Stable network identities such as DNS and hostname.
B ) Stable storage identity.
Before PetSets, stateful applications were supported but exceedingly challenging to deploy and manage, especially regarding distributed stateful clusters. In addition, PODs had random names that could not be relied upon. With the introduction of PetSets controllers and Kubernetes Pets, Kubernetes has sharpened its support for stateful and distributed stateful applications.

Challenges: PODs and their shortcomings
Kubernetes enables the specification of applications as a POD file, expressed in YAML or JSON format. The file specifies what containers are to be in a POD. PODs are the smallest deployment unit in Kubernetes and present several challenges for some stateful services. They do not offer a singleton pattern and are temporary by design.
Their constructs are mortal as they are born and die but never resurrected. When a POD dies, it’s permanently gone and replaced with a new instance and fresh identity. This operation model may suit some applications but falls short for others who want to retain identity and storage across restart / reschedule.
Challenges: Replication Controllers and their shortcomings
If you want a POD to resurrect, you need a Replication Controller ( RC ). The RC enables a singleton pattern to set replication patterns to a specific number of PODs.
The introduction of the RC is undoubtedly a step in the right direction, as it ensures the correct number of replicas are always running at any given time. The RC works alongside services before the RC, using labels to map inbound requests to certain PODs.
Services provide a level of abstraction so that the application endpoint never changes. RC is suitable for application deployments requiring weak uncoupled identities, and when naming individual PODs doesn’t matter to the application architecture.
However, they lack certain functionalities that the new PetSet controller provides. So you could say that a PetSet controller is an enhanced RC controller in shiny new clothes.
A key point: “Pets and Cattle.”
The best way to understand Kubernetes Pets and PetSets is to perceive the cloud infrastructure with the “Pets and Cattle” metaphor. A“Pet” is a special snowflake you have emotional ties towards and requires special handling, for example, when it’s sick, unlike “Cattle,” which is viewed as an easily replaceable commodity.
Cattle are similar enough to each other that you can treat them all as equals. Therefore, the application does not suffer much if a cattle dies or needs replacement.
- Cattle refers to stateless applications, and Pets refer to stateful, “build once, run anywhere” applications.
Discussing Stateless applications
A stateless application takes in a request and responds, but nothing is left behind to fulfill subsequent connections. The stateless pattern derives from another independent system’s ability to satisfy subsequent requests/responses. Stateful applications store data for further use. These types of applications can be inspected with a stateful inspection firewall.
Note: PetSet Objects
Stateful applications are grouped into what’s known as a PetSet object. The PetSet controller has a family-orientated approach, as opposed to the traditional RC, which is mainly concerned with the number of replicas. PODs are stateless disposable units that can be removed and interchanged without affecting the application. The Pets, conversely, are groups of stateful PODs requiring stronger different identities.
Note: PetSet Identities
Within a PetSet, Pets ( stateful applications ) have a unique, distinguishable identity. The identities stick and do not change when restarting/rescheduling. They have an explicit purpose/role in life that is known throughout the family—a definitive startup carried out in a structured order that fits within its responsibility in the application’s framework.
Initially, the cattle approach forced us to view cloud components as anonymous resources. However, this approach does not fit all application requirements. Stateful applications require us to rethink the new pet-style approach.
Workloads and Application Types Suitable for Pets:
Stateful application within a PetSet object requires unique identities such as :
- Storage
- Ordinal index
- Stable Hostname
The PetSet object supports clustered applications that require stricter membership and identity requirements, such as :
- Discovery of peers for quorum
- Startup and Teardown
Workloads that benefit from PetSets include, for example :
- NoSQL databases – clustered software like Cassandra, Zookeeper, etcd, requiring regular membership.
- Relationshional Databases – MySQL or PostgreSQL requiring persistent volumes.
Application Roles & Responsibilities
Applications have different roles and responsibilities, requiring different deployment models. For example, a Cassandra cluster has strict membership and identity requirements; specific nodes are designated seed nodes used during startup to discover the cluster.
They come first and act as the contact points for all other nodes to get information about the cluster. All nodes require one seed node, and all nodes within a cluster must have the same seed node. No node can join the cluster without a seed node, meaning their role is vital for the application framework.
Note: Zookeeper – Identification of Peers
Zookeeper or etcd requires the identification of peers and instances clients should contact. Other databases have a master/slave model where the master has unidirectional control over the slave. The “primary” server has a different role and identity requirements than the “slave.” Properly running these types of services requires more complex features in Kubernetes.
Closing Points on Kubernetes PetSets
PetSets, now more commonly known as StatefulSets, are a Kubernetes resource designed to manage stateful applications. Unlike the standard ReplicaSets that manage stateless applications by treating all replicas as identical, PetSets offer a more sophisticated way to manage pods. Each pod in a PetSet has a guaranteed unique identity, stable networking, and persistent storage, making them ideal for databases, caches, and other applications that require stable identities.
One of the standout features of PetSets is their ability to maintain a stable identity for each pod, which is crucial for stateful applications. This includes:
– **Persistent Storage:** Each pod in a PetSet can have its own persistent volume, ensuring that data is retained across restarts.
– **Ordered Deployment and Scaling:** PetSets ensure that pods are deployed or scaled up in a specific order, which is essential for applications with interdependencies.
– **Stable Network Identity:** Each pod retains its network identity, which is important for applications that rely on consistent access to resources.
Deploying a stateful application using PetSets involves defining a StatefulSet resource in your Kubernetes cluster. This resource specifies the desired number of replicas, the template for the pods, and any associated persistent volumes. By following best practices in your configurations, you can ensure that your stateful applications benefit from the stability and reliability that PetSets offer.
PetSets are particularly beneficial for applications like databases (e.g., MySQL, Cassandra), distributed file systems (e.g., HDFS), and other systems that require stable identities and persistent storage. By utilizing PetSets, organizations can ensure high availability and resilience for their critical stateful applications, even in dynamic environments.
Stateful Failure Investigation
A database member has failed. Investigate the workload before deciding what to do. Stable identity, persistent storage, readiness and quorum all matter.
PVC: db-data-database-0
Role: member
PVC: db-data-database-1
Role: member
PVC: db-data-database-2
Role: member
