Many applications call the same external HTTP endpoints over and over.
That data often changes slowly, or not at all between requests. Yet we still pay the full network and serialization cost on every call.
In .NET we usually reach for HttpClient to make those calls. It works well, but it does not cache responses for you.
That's where Replicant comes in. It wraps HttpClient with disk-based caching, respects standard HTTP cache headers and can make repeated requests much faster.
Replicant
Replicant is a .NET library that adds a caching layer on top of HttpClient.
Instead of storing responses in memory, it writes them to disk. That means cached data survives application restarts and does not compete with your heap.
Replicant follows normal HTTP caching rules. It respects headers like Cache-Control, Expires, ETag and Last-Modified when deciding whether to serve a cached response or revalidate with the server.
When the cache grows beyond a configured limit, older entries are removed based on last access time.
You can use it in two main ways:
- HttpCache - A standalone wrapper that owns an HttpClient instance
- ReplicantHandler - A DelegatingHandler that plugs into an existing HttpClient pipeline
Getting Started
To get started, install the NuGet package. You can do this via the NuGet Package Manager or by running the following command in the Package Manager Console:
Install-Package Replicant
For the examples in this post I used version 2.0.3.
HttpCache
The simplest way to try Replicant is through HttpCache.
Replicant ships with a default static instance you can use right away:
var content = await HttpCache.Default.DownloadAsync("https://httpbin.org/status/200");
This caches responses under {Temp}/Replicant.
For real applications you should create a long-running instance and dispose it when the application shuts down:
await using var httpCache = new HttpCache(
cacheDirectory,
new HttpClient
{
Timeout = TimeSpan.FromSeconds(30)
},
maxEntries: 10000);
var content = await httpCache.StringAsync("https://httpbin.org/json");
HttpCache exposes several helpers depending on what you need from the response:
- StringAsync - Returns the response body as a string
- BytesAsync - Returns the response body as a byte array
- StreamAsync - Returns the response body as a stream
- ResponseAsync - Returns the full HttpResponseMessage
When using dependency injection, register HttpCache as a singleton:
builder.Services.AddHttpClient();
builder.Services.AddSingleton(_ =>
{
var clientFactory = _.GetRequiredService<IHttpClientFactory>();
return new HttpCache(cacheDirectory, clientFactory.CreateClient);
});
Pairing HttpCache with IHttpClientFactory avoids the socket exhaustion issues that come from creating a new HttpClient for every request.
Sample Setup
To show the difference in practice, I put together a small solution with two projects.
The first project is a Web API backed by PostgreSQL. It exposes a single endpoint that returns the latest 1000 products from the database:
app.MapGet("/products", async (ApplicationDbContext dbContext, CancellationToken cancellationToken) =>
{
var products = await dbContext.Products
.OrderBy(x => x.CreatedAt)
.Take(1000)
.ToListAsync(cancellationToken);
return Results.Ok(products);
});
The database is seeded with one million products so each response carries a meaningful payload.
Everything runs through Docker Compose. Start the stack, apply migrations and run the seed script from the sample project.
The second project is a simple console app. It calls the same endpoint five times with a plain HttpClient and five times with Replicant, then prints the elapsed time for each request.
Comparing Performance
Here is the HttpClient version from the consumer app:
async Task WithHttpClient()
{
var httpClient = new HttpClient();
for (int i = 0; i < 5; i++)
{
var stopwatch = Stopwatch.StartNew();
var response = await httpClient.GetAsync("https://localhost:5001/products");
stopwatch.Stop();
Console.WriteLine($"Elapsed time: {stopwatch.Elapsed.TotalMilliseconds} ms");
}
}
Every iteration hits the network and waits for the full response.
The Replicant version uses HttpCache.Default and the same URL:
async Task WithReplicant()
{
var httpClient = HttpCache.Default;
for (int i = 0; i < 5; i++)
{
var stopwatch = Stopwatch.StartNew();
var response = await httpClient.ResponseAsync("https://localhost:5001/products");
stopwatch.Stop();
Console.WriteLine($"Elapsed time: {stopwatch.Elapsed.TotalMilliseconds} ms");
}
}
The first call behaves like a normal HttpClient request. After that, Replicant serves the cached response from disk and the elapsed time drops sharply.
On my machine the HttpClient runs stayed around several seconds each. With Replicant, the first request was similar and the following four were near instant.
Exact numbers will depend on your hardware and network, but the pattern is consistent. Disk cache hits avoid the repeated download cost.
ReplicantHandler
If you already have an HttpClient pipeline, ReplicantHandler is the better fit.
It works like any other delegating handler. You can wire it up manually or through HttpClientFactory:
builder.Services.AddReplicantCache(cacheDirectory);
builder.Services.AddHttpClient("CachedClient")
.AddReplicantCaching();
Only GET and HEAD requests go through the cache. POST, PUT and DELETE pass straight through to the server.
Replicant also supports a few useful options worth knowing about:
- staleIfError - Serve the last good cached response when revalidation fails
- maxRetries - Retry transient HTTP failures with exponential backoff
- minFreshness - Keep entries fresh for a minimum duration even when the server sends short expiry headers
Replicant can even act as a disk-based L2 backend for HybridCache through ReplicantDistributedCache.
For full API details and advanced scenarios like composing with resilience pipelines, check the official Replicant GitHub repository.
Conclusion
Repeated HTTP calls to the same endpoint are a common performance drain.
Replicant gives you a straightforward disk-based caching layer on top of HttpClient. It respects standard cache headers, survives restarts and fits into both standalone apps and existing HttpClientFactory pipelines.
For read-heavy workloads where data changes slowly, it is a simple way to cut response times without building your own cache from scratch.
If you want to check out examples I created, you can find the source code here:
Source CodeI hope you enjoyed it, subscribe and get a notification when a new blog is up!



