---
title: Hiya’s best practices around Kafka consistency and availability
description: Apache Kafka’s distributed, durable, and high-throughput nature makes it a natural fit for streaming many types of data. Hiya uses Kafka for a number of critical use cases, such as asynchronous data processing, cross-region replication, storing service logs, and more
image: https://blog.hiya.com/hubfs/image-620x420.png
---

[![Hiya logo](https://blog.hiya.com/hubfs/hiya-new-purple.svg)](https://www.hiya.com/)

Solutions

Solutions

[For Mobile Operators Protect your network](https://www.hiya.com/solutions/operators) [For Businesses Brand your calls](https://www.hiya.com/solutions/businesses) [For People Get the mobile app](https://www.hiya.com/products/apps/hiya-spam-blocker)

###### Webinar

Branded Calling 101 Weekly Webinar

Make your company's calls more recognizable. Learn how Hiya can drive value for your business.

[Sign up today Arrow Right](https://www.hiya.com/lp/webinar-series-intro-to-hiya-branded-call)

Products

 Connect

###### [Branded Call Display your branded caller ID](https://www.hiya.com/products/connect/branded-call)

###### [Number Registration Display your branded caller ID](https://www.hiya.com/products/connect/number-registration)

###### [View Plans Flexible pricing for teams of all sizes](https://www.hiya.com/products/connect/pricing)

 Protect

###### [Spam Analytics Stop spam & fraud on your mobile network](https://www.hiya.com/products/protect/spam-analytics)

###### [AI Voice Detection Real-time AI voice detection](https://www.hiya.com/products/protect/ai-voice-detection)

 Mobile apps

###### [Hiya Spam Blocker Fraud & AI voice protection](https://www.hiya.com/products/apps/hiya-spam-blocker)

###### [Hiya AI Phone Productivity for busy people](https://www.hiya.com/products/apps/hiya-ai-phone)

![](https://blog.hiya.com/hubfs/sotc-report.avif)

###### Report

State of the Call 2026

86% of unidentified calls go unanswered. Read the benchmark report for what is happening in voice today, and what you can do to drive business.

[Read the report Arrow Right](https://www.hiya.com/state-of-the-call)

Why Hiya?

Overview

[Why Hiya Your voice innovation partner](https://www.hiya.com/why-hiya) [How it works Get started quickly & easily](https://www.hiya.com/how-it-works) [Customer Stories Real companies, real results](https://www.hiya.com/case-studies) [Voice Intelligence Platform Industry’s leading voice platform](https://www.hiya.com/why-hiya/voice-intelligence-platform) [Trust Center Compliance, security, & privacy](https://www.hiya.com/why-hiya/trust-center)

Company

[About Leadership and history](https://www.hiya.com/company/about) [Careers We're hiring!](https://www.hiya.com/company/careers) [Contact us Get in touch](https://www.hiya.com/contact-us)

![](https://blog.hiya.com/hubfs/bclc-hero-img.webp)

###### Customer Story

BCLC increases business KPIs with Hiya

With Hiya Branded Call BCLC was able to increase contact rates, campaign efficiency, and revenue.

[Read their story Arrow Right](https://www.hiya.com/customer-stories/bclc)

Resources

Resources

[Resource Center](https://www.hiya.com/resources) [Partner Program](https://partners.hiya.com/) [Get Support](https://hiya.com/support) [Developer Docs](https://developer.hiya.com/)

[Hiya Blog](https://blog.hiya.com/) [Newsroom](https://www.hiya.com/newsroom) [Events](https://www.hiya.com/events)

###### eBook

10 Tips for customer-friendly phone calls

 Prevent caller reputation issues and complaints with customer-friendly calling practices. 

[Read eBook Arrow Right](https://www.hiya.com/resources/ebooks/10-tips-for-customer-friendly-phone-calls)

[Get Started Arrow Right](https://www.hiya.com/#audience)

[Log in](https://connect.hiya.com/login) [Get started](https://www.hiya.com/#audience)

- [Home](https://hiya.com)
- [Blog](https://blog.hiya.com)
- [Hiya’s best practices around Kafka consistency and ...](https://blog.hiya.com/hiyas-best-practices-around-kafka-consistency-and-availability/)

# Hiya’s best practices around Kafka consistency and availability

[Jake Utley](https://blog.hiya.com/author/jutley)

Sep. 19, 2020

[![](https://blog.hiya.com/hubfs/icon-fb.svg)](https://www.facebook.com/sharer/sharer.php?u=https://blog.hiya.com/hiyas-best-practices-around-kafka-consistency-and-availability/) [![](https://blog.hiya.com/hubfs/icon-x.svg)](https://twitter.com/intent/tweet?url=https://blog.hiya.com/hiyas-best-practices-around-kafka-consistency-and-availability/&text=Hiya%E2%80%99s+best+practices+around+Kafka+consistency+and+availability) [![](https://blog.hiya.com/hubfs/icon-linkedin.svg)](https://www.linkedin.com/shareArticle?mini=true&url=https://blog.hiya.com/hiyas-best-practices-around-kafka-consistency-and-availability/)

## At a glance

![](https://blog.hiya.com/hubfs/image-620x420.png)

## Introduction

Apache Kafka’s distributed, durable, and high-throughput nature makes it a natural fit for streaming many types of data. Hiya uses Kafka for a number of critical use cases, such as asynchronous data processing, cross-region replication, storing service logs, and more. Since Kafka is a central component of so many pipelines, it’s crucial that we use it in a way that ensures message delivery.

In this article I’ll share some of our best practices for ensuring consistency and maximizing availability when using Kafka. This material is based off a brown bag presentation I gave for our engineering org on the same topic. The slides for the original presentation can be found [here](https://docs.google.com/presentation/d/16Kr2WJSJ00TUMp_Vbl0IWCjzLI6Cifu9e2AomG80KR8/edit?usp=sharing).

## Necessary Background

Before we continue, let’s review some of the fundamentals of Kafka.

When writing messages into Kafka, we write them into a *topic*. Topics contain logically related messages which are divided into a number of *partitions*. Partitions serve as our unit or ordering, replication, and parallelism. While topics are the interesting component from a business perspective, operationally speaking, partitions are the star of the show.

Topics are configured with a *replication-factor*, which determines the number of copies of each partition we have. All replicas of a partition exist on separate brokers (the nodes of the Kafka cluster). This means that we cannot have more replicas of a partition than we have nodes in the cluster.

A replica is either the leader of its partition, or a follower of the leader. A follower can either be in-sync with the leader (contains all the partition-leader’s messages, except for messages within a small buffer window), or out-of-sync. The set of all in-sync replicas (including the partition-leader) is referred to as the ISR.

This information is summarized by the following illustration of a topic with a partition-count of 3 and a replication-factor of 2:

 

![0_o0afnsjKw5sb2Z9a](https://blog.hiya.com/hs-fs/hubfs/0_o0afnsjKw5sb2Z9a.png?width=600&name=0_o0afnsjKw5sb2Z9a.png)

 

A topic with 3 partitions and 2 replicas spread across 3 brokers.

## Understanding Broker Outages

In order to properly tolerate a broker outage, we must first understand the behavior of a Kafka cluster during a broker outage. Let’s explore this in enough depth to address our concerns around partition availability and consistency.

During a broker outage, all partition replicas on the broker become unavailable, so the affected partitions’ availability is determined by the existence and status of their other replicas. If a partition has no additional replicas, the partition becomes unavailable. If a partition has additional replicas that are in-sync, one of these in-sync replicas will become the interim partition leader. Finally, if the partition has addition replicas but none are in-sync, we have a choice to make: either we choose to wait for the partition leader to come back online–sacrificing availability — or allow an out-of-sync replica to become the interim partition leader–sacrificing consistency.

Once the broker comes back online, all its replicas will come back in sync with their partition-leaders. Once caught up, any replicas that were originally partition-leaders will become partition-leaders once again. It is worth noting that during the catch-up phase, the replicas are checking for missing messages based solely on **message offsets**. This detail will become important later on.

## Handling Broker Outages

### Replication-Factor and Minimum ISR Size

Based on this behavior, we can see that a partition must have an extra in-sync replica available to survive the loss of the partition leader. In order to satisfy this requirements, we can configure our topic’s \`min.insync.replicas\` setting to 2. If we want all topics to have this setting by default, we can configure our brokers with an identical setting.

In order to have a minimum ISR size of 2, our replication-factor must be at least this value. However, we actually need to be more strict. If we only have 2 replicas and then lose a broker, our ISR size shrinks to 1–below our desired minimum. Therefore, we must set our replication-factor to be greater than the minimum ISR size (at least 3). To make this the topic default, we can configure our brokers by setting \`default.replication.factor\` to 3.

### Acks

With these settings, our topics are configured to tolerate a single broker failure, but our producers still have a role to play to ensure consistency. When a producer sends a produce request, the conditions for the request to be considered a success is determined by a producer configuration parameter called \`acks\`, which can be set to one of the following values:

- 0: return success at the moment the request is sent
- 1: return success once the partition leader acknowledges the request
- all: return success once all replicas in the ISR acknowledge the request

Consider this situation where this setting is set to either 0 or 1. The producer sends a message X to a partition, then the partition leader dies shortly after. In this situation, X will get written to one replica, but not the others. Despite this, our other replicas are still in-sync (messages within the buffer window are not required to be in-sync) and one will become the interim partition leader. Now, when the producer sends a messages Y to the partition, it will be written *with the same offset* as message X. Once the down broker comes back online, the partition leader will not realize that it is missing message Y, and our partition has become inconsistent.

If we had set \`acks\` to \`all\` instead, then the request with message X would have either been written to all our replicas or the request would have failed. In either case we avoid the demonstrated inconsistency.

 

![0_XHm-1o52h-BLDQp6](https://blog.hiya.com/hs-fs/hubfs/0_XHm-1o52h-BLDQp6.png?width=579&name=0_XHm-1o52h-BLDQp6.png)

Setting acks to 0 or 1 can lead to inconsistent partitions

## Duplicates

While the specified settings do give us consistency with high availability, there is still a detail that needs to be addressed: these settings provide us with at-least-once message delivery. On other words, we will end up with some duplicate messages in our topics. To see this, consider the following scenario.

Using the recommended configuration, our producer writes a message, which gets written to the partition leader. Before the follower replicas fetch the message, the request times out. Shortly after, the follower replicas fetch the message. In this situation, the message was written successfully to all replicas, but the producer considers the request to have failed, and retries. After a successful retry, we’ve written the message to our partition **twice**.

Thankfully, at-least-once message delivery does not have to be a problem. It does, however, require consideration in how we design our consumers. In particular, consumers should be designed to handle messages [*idempotently*](https://en.wikipedia.org/wiki/Idempotence). This can be accomplished in a variety of ways, such as by including a unique id to each message, then tracking the processed ids. If we come across a message with an id that we’ve already processed, we simply skip over the message.

If you would rather design around exactly-once message delivery, the good news is that Kafka does support it! However, it is not straightforward, so I will not explain it in this article. Instead, consult [this article](https://www.confluent.io/blog/exactly-once-semantics-are-possible-heres-how-apache-kafka-does-it/) from Confluent’s Neha Narkhede.

## AZ outages

Based on what we’ve learned so far, we can ensure that even if a broker goes down, we can ensure that all partitions of a topic remain available without sacrificing consistency. This will save us from many failures, but not all! We still need to be concerned about the possibility of an entire datacenter outage.

Thankfully Kafka makes it very easy to handle datacenter outages using a feature called *Rack Awareness*. The idea is simple: each broker is configured with a label describing which “rack” (or datacenter) the broker is within. Then, when Kafka assigns replicas across different brokers, it spreads our replicas across the available racks. If we have 3 Kafka brokers spread across 3 datacenters, then a partition with 3 replicas will never have multiple replicas in the same datacenter. With this configuration, datacenter outages are not significantly different from broker outages.

## Concurrent failures

Of course, we can always experience a combination of failure modes. What happens when multiple brokers fail in different AZs? What happens when multiple AZs fail? I will not dive into these questions, but if they intrigue you, consider them an exercise to the reader.

## A brief note on Zookeeper

While these settings can help ensure consistency and high uptime in your Kafka topics, remember that Kafka is dependent on Zookeeper. Therefore, a Kafka cluster is *at most* as reliable as the Zookeeper it depends on. Make sure to put some thought into how your Zookeeper is configured to reach your HA requirements.

[![banner-hiya-reg](https://blog.hiya.com/hubfs/banner-hiya-reg.jpg)](https://www.hiya.com/products/connect/number-registration)

### Latest Articles ⚡️

- <https://blog.hiya.com/hiya-awarded-16-g2-summer-2026-badges-for-branded-caller-id?hsLang=en>
  
   Hiya awarded 16 G2 Summer 2026 badges for branded caller ID
  
  Lena Prickett
  
  Jul. 9, 2026
- <https://blog.hiya.com/how-canada-is-leading-the-fight-to-stop-bank-scams?hsLang=en>
  
   How Canada is leading the fight to stop bank scams
  
  Stephanie Boulanger
  
  Jun. 26, 2026
- <https://blog.hiya.com/how-to-check-phone-numbers-for-spam-labels?hsLang=en>
  
   How to check phone numbers for spam labels
  
  Michelle Wallace
  
  Jun. 9, 2026

## Related articles

<https://blog.hiya.com/how-canada-is-leading-the-fight-to-stop-bank-scams?hsLang=en> ![](https://blog.hiya.com/hubfs/rogers-toronto-event-blog-hero.webp)

## [How Canada is leading the fight to stop bank scams](https://blog.hiya.com/how-canada-is-leading-the-fight-to-stop-bank-scams?hsLang=en)

Stephanie Boulanger

Jun. 26, 2026

<https://blog.hiya.com/how-to-check-phone-numbers-for-spam-labels?hsLang=en> ![](https://blog.hiya.com/hubfs/Blog_How-to-check-phone-numbers-for-spam-labels_B.webp)

## [How to check phone numbers for spam labels](https://blog.hiya.com/how-to-check-phone-numbers-for-spam-labels?hsLang=en)

Michelle Wallace

Jun. 9, 2026

<https://blog.hiya.com/hiya-and-dna-finland-launch-network-level-call-protection?hsLang=en> ![](https://blog.hiya.com/hubfs/Hiya%20x%20DNA-blog@2x.webp)

## [Hiya and DNA Finland launch network-level call protection](https://blog.hiya.com/hiya-and-dna-finland-launch-network-level-call-protection?hsLang=en)

Stephanie Boulanger

Jun. 2, 2026

[![](https://blog.hiya.com/hubfs/hiya-new-purple.svg)](https://www.hiya.com/)

[![](https://blog.hiya.com/hubfs/linkedin-1.svg)](https://www.linkedin.com/company/hiyainc) [![](https://blog.hiya.com/hubfs/facebook-1.svg)](https://www.facebook.com/hiyainc/)

###### Products

- Connect
- [Number Registration](https://www.hiya.com/products/connect/number-registration)
- [Branded Call](https://www.hiya.com/products/connect/branded-call)
- [View Plans](https://www.hiya.com/products/connect/pricing)
- Protect
- [Spam Analytics](https://www.hiya.com/products/protect/spam-analytics)
- [AI Voice Detection](https://www.hiya.com/products/protect/ai-voice-detection)
- Apps
- [Hiya Spam Blocker](https://www.hiya.com/products/apps/hiya-spam-blocker)
- [Hiya AI Phone](https://www.hiya.com/products/apps/hiya-ai-phone)

###### Solutions

- Company size
- [Enterprise](https://www.hiya.com/solutions/enterprise)
- [Call Centers](https://www.hiya.com/solutions/call-center)
- [Small and Medium](https://www.hiya.com/solutions/businesses)
- Service providers
- [Operators](https://www.hiya.com/solutions/operators)
- [OEMs and technology](https://www.hiya.com/solutions/technology-partners)

###### Considering Hiya?

- [Why Hiya](https://www.hiya.com/why-hiya)
- [How it Works](https://www.hiya.com/how-it-works)
- [Customer Stories](https://www.hiya.com/case-studies)
- [Voice Intelligence Platform](https://www.hiya.com/why-hiya/voice-intelligence-platform)
- [Trust Center](https://www.hiya.com/why-hiya/trust-center)
- [Modern Slavery](https://www.hiya.com/company/modern-slavery)
- [About Hiya](https://www.hiya.com/company/about)
- [Careers: We're hiring!](https://www.hiya.com/company/careers)
- [Contact us](https://www.hiya.com/contact-us)

###### Resources

- [Resource Center](https://www.hiya.com/resources)
- [Partner Program](https://partners.hiya.com/)
- [Get Support](https://hiya.com/support)
- [Developer Docs](https://developer.hiya.com/)
- [Hiya Blog](https://hiya.com/blog)
- [Events](https://www.hiya.com/events)
- [Press Kit](https://www.hiya.com/newsroom#press-kit)

![ANAB](https://blog.hiya.com/hubfs/logo-anab.svg) ![ISO 27001](https://blog.hiya.com/hubfs/logo-iso27017.png) ![AICPA](https://blog.hiya.com/hubfs/logo-aicpa.svg) ![IAF](https://blog.hiya.com/hubfs/logo-iaf.svg)

[Privacy Policy](https://www.hiya.com/legal/privacy) [Terms of Service](https://www.hiya.com/legal/terms-of-service) [App Data Privacy Policy](https://www.hiya.com/legal/app-data-protection-and-privacy-policy)

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Jake Utley",
    "url" : "https://blog.hiya.com/author/jutley"
  },
  "dateModified" : "2020-09-29T23:09:39.021Z",
  "datePublished" : "2020-09-19T23:13:10.000Z",
  "headline" : "Hiya’s best practices around Kafka consistency and availability",
  "image" : [ "https://blog.hiya.com/hubfs/image-620x420.png" ],
  "mainEntityOfPage" : {
    "@id" : "https://blog.hiya.com/hiyas-best-practices-around-kafka-consistency-and-availability/",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://blog.hiya.com/hubfs/hiya-logo-3.svg"
    },
    "name" : "Hiya Inc."
  }
}
```