Several years ago when I was interviewing at Calavista, my future colleagues asked me what I was most proud of in my career. At the time I was mostly known for my writing years earlier at CodeBetter and StructureMap, but even then my obvious answer was Marten. Wolverine was even then my true passion project, but Marten is undeniably the biggest technical success I’ve ever had and the single biggest source of JasperFx Software business today.
Sometime this past weekend, Marten quietly rolled past 20 million downloads on NuGet. Almost exactly eleven years after the first commit in October of 2015, the “PostgreSQL as a document database” experiment that started as a tactical way to get off of RavenDB in a hurry has become the most widely used and, I’ll very happily argue, the most capable Event Sourcing solution in the .NET ecosystem.
Don’t write off Marten because it’s a FOSS tool in the .NET ecosystem — the perennial Harvey Dangerfield of technical platforms. Marten, especially when paired up with Wolverine, has a feature set that I believe is very competitive and in many cases superior to the commercial event sourcing tools in the JVM.
I wrote about Marten’s 10th birthday last year and a five year retrospective on adoption this spring, so I’ll keep this short. Twenty million is a nice round number though, and it’s worth a moment to say thank you and take stock.
The numbers
Straight from GitHub and NuGet this morning:
NuGet downloads
20,077,707
GitHub stars
3,460
Contributors
257
Pull requests merged
2,275
Issues closed
2,594 of 2,605 ever opened
Tagged releases
321
Latest release
V9.41.0, shipped today
Two of those matter more to me than the download count. 257 contributors is a real community (that’s a random Wednesday for a popular npm project, but huge for .NET), not a one-person project with a mailing list. And eleven open issues on a project this old and this large is the result of a lot of people caring about a lot of details for a very long time. Thank you to everyone who has filed an issue, sent a pull request, answered a question in Discord, or written up their experience for others.
For context on “most widely used”: the next closest event store package on NuGet, the official client for what used to be called EventStoreDB, sits at a bit over 8 million downloads. SqlStreamStore, EventFlow, and Eventuous are each in the low single-digit millions or below. Marten isn’t just in front. It’s in front by a wide margin, and the gap is widening.
Marten as a document database
Plenty of those downloads aren’t for Event Sourcing at all. Marten’s original pitch was simple: PostgreSQL is a fantastic database, JSONB is a fantastic storage format, and .NET developers deserve a document database experience without giving up transactions, a real query engine, or the operational know-how they already have. That pitch still holds.
You’ll see plenty of people online try to say that you can’t really use a relational database as the foundation for an event store, but our PostgreSQL foundation means that Marten can support a much stronger set of projection capabilities and genuine options for strong consistency that the “dedicated event store databases” cannot.
Marten is the most complete Event Sourcing toolkit on .NET, and it isn’t particularly close. Most event store products give you an append-only log and a subscription API, then wish you luck with everything else. Marten’s position has always been that the event store is only useful if the projections, read models, consistency model, and operational story are all in the box. Highlights that I think are unique, or at least unusually strong, in the .NET world:
Events and read models live in one database. You get transactional consistency between writes and inline projections, query read models with the same LINQ provider, and rebuilding a projection is one command instead of a migration project.
Dynamic Consistency Boundary. Marten 9 shipped first class DCB support for enforcing invariants across multiple entities without one giant aggregate stream. As far as I know it’s still the only mainstream .NET event store where this is built in. See Higher Performance DCB Development with Marten 9.0.
Wolverine. The Wolverine integration turns Marten into a full CQRS platform: the aggregate handler workflow (our version of the “Decider” pattern), a transactional outbox on the same database, event subscriptions feeding handlers, and HTTP endpoints that stream JSON straight out of PostgreSQL. That combination is what people mean by “the Critter Stack.”
Marten’s model has also become the template for the rest of the stack. Polecat brings the same projections and aggregates to SQL Server 2025, and Fisher does the same for embedded scenarios. Learn Marten and you’ve learned all three.
Where JasperFx fits
Marten is and will stay MIT licensed open source. What’s changed is that there’s now a real company behind it. JasperFx Software exists to make the Critter Stack sustainable for the long haul, and I explained the model in The “Open Core” Model for Sustainable OSS Development in .NET. If your team depends on Marten in production:
CritterWatch is our monitoring and operations console for Marten, Polecat, and Wolverine, with visibility into projections, subscriptions, dead letters, and message flow across a whole fleet, plus an MCP server for AI assisted production support. Plans are on the products page.
Critter Stack AI Skills teach Claude Code and other coding agents to write idiomatic Marten, Wolverine, and Polecat code the way we would. The latest additions are in AI Skills 1.10.
Thank you
Twenty million downloads, 257 contributors, eleven years. Marten got here because a lot of people decided PostgreSQL plus .NET was a good idea and kept showing up to make it better. If you’ve never tried it, the quick start will have you appending events and building a projection in about ten minutes. If you have, thank you. Here’s to the next 20 million.
Critter Stack Roundup: Two Weeks, Six Repositories, and an EF Core Sweep
It’s been a busy couple of weeks across the Critter Stack. Between September 13th and today, we merged roughly 370 pull requests across JasperFx, Weasel, Marten, Polecat, Fisher, and Wolverine, and shipped a pile of releases along the way. More importantly, a big chunk of that work came from, or was driven by, people outside of JasperFx Software, so there’s a lot of thanking to do in this post.
5.27.0, 5.28.0, 5.28.1, 5.29.0, 5.30.0, 5.31.0, and now 5.32.0
Weasel
9.33.0, 9.34.0, 9.35.0, 9.35.1, 9.35.2
JasperFx
2.69.1 through 2.75.2
Fisher
1.5.0 through 1.13.0
And Polecat 5.32 will be released by the time you read this.
And the big themes:
An Entity Framework Core sweep through Wolverine and Weasel, touched off by a set of sharp bug reports — I wouldn’t expect any of the bugs we address to impact many people, but I’d like the number of EF Core features not supported or features we don’t support well to be essentially zero
Vector, full text, and hybrid search across Marten, Polecat, and Fisher, all behind one shared API – this is to support our forthcoming agentic memory tool “Stoat”
Conjoined multi-tenancy tightened up and held to a shared compliance suite on every store. This was mostly a testing effort, but there were some EF Core related issues that helped spawn that work. It did find some issues with Polecat and Fisher when we did more compliance testing against Marten behavior to the younger tools
Error messages that tell you what to do instead of just telling you something went wrong – I wrote about that previously in our new AI Skills 1.14 release notes
Idempotency and agent distribution improvements in Wolverine, a lot of it from the community – Some things are just flat out hard. Gulp.
The EF Core Sweep
This one started with a trio of excellent issue reports from Ali Yuksekkaya against Wolverine’s EF Core integration:
Conjoined EF Core tenancy was ignored in the Lightweight transaction mode, so a message could be written under the wrong tenant and HTTP cascades skipped the outbox (GH-4611)
Conjoined tenancy allowed detached updates and deletes to touch rows that belonged to another tenant (GH-4612)
Returning Update<T> or Store<T> from a handler saved nothing at all for an entity the DbContext wasn’t already tracking (GH-4613)
Those got fixed, but reports like that are a signal that nobody had been looking hard enough at that corner of the code, so we went looking. The resulting sweep landed as a wave of Wolverine pull requests this weekend:
Lightweight EF Core message handlers now get a real transactional outbox. Before, Lightweight mode quietly meant no outbox enlistment, no domain event scraping, and no idempotency check
Domain event envelopes scraped out of the DbContext are now flushed before the Eager transaction commits. This was a leftover from an earlier fix where the scraped envelopes were tracked after the only SaveChanges() call and never persisted
Wolverine’s conjoined tenant query filter now composes with your query filter instead of replacing it. EF Core 9 and EF Core 10 behave differently here (EF 9 discards the earlier filter, EF 10 throws), and we now cover both
A new EfCoreOps family of declarative side effects for ExecuteUpdate, ExecuteDelete, raw SQL, and bulk inserts. These force the Eager transaction mode they need and scope themselves to the current tenant
A conjoined tenancy test battery that now runs against Marten, Polecat, and Fisher backed message stores
Closing several test coverage holes, including owned, complex, and JSON mapped models end to end
::: warning One of these changes is technically breaking. If you’re using AutoApplyTransactions() and a single handler could be claimed by two persistence providers (say, it takes both a DbContext and a Marten IDocumentSession), Wolverine used to silently apply no transaction at all. It now fails loudly at startup and tells you how to resolve it (GH-4631). If your application hits this, it had a real bug that this change is surfacing. :::
The sweep reached down into Weasel as well, which is where our EF Core schema migration support lives. Ali also reported that EF Core batched queries silently returned incomplete entities for owned, complex, and JSON members, and then contributed a follow up pull request to prepare each batched query once while keeping the provider’s parameter types intact. On top of that, Weasel now:
Materializes EF Core batched queries through EF Core itself
Carries database indexes that EF Core can’t model as their own DDL
Maps the columns of table-split complex properties, which were previously omitted and then dropped by CreateOrUpdate
Never drops a column from an EF Core derived table just because the model doesn’t happen to declare it
Marten, Polecat, and Fisher also all fixed the same bug where an EF Core backed projection leaked the DbContext it created per batch. Thanks to wpei-infotrack for reporting the Marten version of that one, which turned out to be a leaked PostgreSQL connection per batch in the async daemon.
Vector, Full Text, and Hybrid Search Everywhere
The other big feature push was around search. JasperFx now has a shared, store-neutral vector and hybrid search surface in JasperFx.Events.Vectors, and all three of our document stores implement it:
Marten.PgVector moved onto the shared contracts with scored search, HNSW index declarations, and hybrid search using reciprocal rank fusion over PostgreSQL’s ts_rank and the vector leg. Marten also now warns you when a full text search falls back to an unindexed, whole document scan
Polecat picked up vector search on SQL Server 2025’s native VECTOR type, a Polecat-owned full text inverted index with LINQ operators and BM25 scoring, prefix search, and hybrid search on top of both
Fisher got hybrid search and embeddings produced from an event stream
Because they all implement the same contract, there’s now a DocumentSearchCompliance suite in JasperFx that holds all three stores to the same behavior. As usual, the first real run of that suite found defects in the suite itself as well as in the stores, which is exactly what it’s for.
Polecat 5.32
Polecat 5.32 is the release that rolls up the last few days of work, and it’s largely about parity with Marten and about multi-tenancy:
Conjoined document tenancy sweep. Every document shape is now tested in both directions, and conjoined tenancy now also holds on the event, vector, and partition onboarding paths
Raw SQL in IBatchedQuery, bringing batching up to parity with Marten
Strongly typed identifiers are now assigned onto a live aggregated aggregate, just like Marten does
Document indexes, computed columns, and foreign keys are now modeled as Weasel schema objects, which means db-dump finally reproduces the full configured schema
Every tenant database is described in the store’s usage descriptor for a multi-tenanted store, which matters for CritterWatch
Event store diagnostic reads answer “no results” rather than throwing when the schema was never applied or has drifted
Adoption of JasperFx 2.75 and Weasel 9.35
Earlier in the window, Polecat 5.31 also made startup migrations take a real cross-process lock through sp_getapplock and routed the last few hand-escaped SQL construction sites through a shared escaping helper.
Multi-Tenancy, Everywhere
Multi-tenancy was a recurring thread through all six repositories:
JasperFx now has conjoined document tenancy compliance tests, and Marten, Polecat, and Fisher all enrolled
TenantIdStyle is now applied consistently. Marten applies it at every boundary that stamped or keyed on the raw tenant id, and Wolverine now normalizes Envelope.TenantId through it so that the stores reading that value write the right tenant
There’s a new canonical DisabledTenantException in JasperFx. All the stores and Wolverine now refuse a disabled tenant with that exception instead of reporting “Unknown tenant id”
IEventStore.OpenReadOnlyEventStore(tenantId) is now tenant aware, so the read-only tier is actually reachable in multi-tenanted systems
Marten and Fisher both fixed bugs with the diagnostic and explorer reads that CritterWatch depends on, including one in Marten where a read against an unknown tenant under sharded tenancy could provision a new tenant and run DDL
Error Messages That Name the Remedy
I spent a chunk of the last two weeks going through the exception messages across the stack, asking one question of each: does this tell the user what to do next? A lot of them didn’t. That turned into a wave of small pull requests:
Wolverine saga failures, handler discovery, missing aggregates, oversized Azure Service Bus messages, mismatched RabbitMQ queue declarations, missing Redis streams, missing HTTP transport clients, and SNS configuration problems all name the remedy now. Named connection strings are validated in one pass at startup. The SignalR transport fails the host start if the hub refuses the connection. Wolverine.HTTP gets a one-line opt in for mapping concurrency failures to a 409 ProblemDetails, and an unknown tenant id maps to a 404 instead of a 500
Marten stream identity mismatches, stream collisions, LINQ refusals, and the rich append concurrency exception all got clearer
Weasel decodes sp_getapplock failures, translates database permission failures into a typed exception, and now warns before AutoCreate.All drops and recreates an object
JasperFx lifted canonical ArchivedStreamException, DisabledTenantException, and stream exceptions so that all three stores throw the same types with the same guidance
On a related note, Marten 9.39 includes two SQL injection fixes, for GroupBy()HAVING comparison operands and for full text search regConfig values on every sink, not just the WHERE clause. If you’re on an older 9.x version, please upgrade.
Wolverine
Besides the EF Core work above, here are some of the highlights in Wolverine:
Capacity aware agent assignment.Michael Harris contributed per-node capacity ceilings for agent distribution (GH-3959), so one node dying no longer pushes its entire share onto the survivors. Anne Erdtsieck filed the original issue, and also contributed a fix for group affinity placement during blue/green deployments. There’s new documentation for the whole thing
Transactional deduplication. Wolverine’s deduplication claim can now ride the native Marten, Polecat, or Fisher transaction. Laurence Gillian reported that an HTTP deduplication claim survived a non-2xx response and turned legitimate retries into false duplicates, and that’s fixed too
An Oracle external table transport contributed by Travis Kirke, along with a fix for the Oracle durability agent’s incoming message recovery
GCP Pub/Sub now shares one subscription across nodes by default, thanks to a report from bittercoder about duplicated messages
Topology scoped message grouping rules, from a request by Anne Erdtsieck
Recurring schedule operability with occurrence attribution and a manual trigger, and a fix for non-UTC recurring schedules. Both came from issues filed by Babu Annamalai
OpenTelemetry parenting fixes. Recurring messages, inline receivers, and Wolverine’s internal agent loops no longer inherit whatever Activity happened to be current when they were started. Marten had a similar fix for spans being re-parented to their grandparent, reported by bohdan-hukivskyi. Open Telemetry sometimes has some weird behavior in terms of how parents are tracked. I expect or hope this will help the CritterWatch graphing of Otel spans from Wolverine
Tore Hammervoll fixed TypeLoadMode.Static so a handler chain finds its pre-generated type by full name instead of scanning exported types per chain
Two concurrency fixes reported by Marcin Aumiler: delayed sends to a partitioned PostgreSQL queue could be deleted without ever being handled, and the listener collection could be corrupted when agents started in parallel
Marten
Other than the search work, multi-tenancy, and messages, Marten had a lot of community driven fixes:
Anne Erdtsieck fixed the outer projection of GroupJoin/SelectMany and GROUP BY rendering over a join, made the projection batch fault properly when an operation can’t be configured, and made an unprovisioned event store answer “nothing” for its progression and dead letter tables
Erik Shafer fixed patched documents and replaced events to be stamped with the session’s actual instant (reported by BaerMitUmlaut)
vpetrevski routed QuickWithServerTimestamps stream starts through mt_quick_append_events to avoid sequence gaps
Arnel Robles corrected the pgvector docs and reported two async daemon bugs in the skip-ahead loader and progression writes
tychomensing-topicus reported a Select() projection problem with absent JSON keys
Weasel
Besides the EF Core work, Weasel got two nice community contributions. Joel Reinford made SQL Server migration scripts runnable under sqlcmd and safe to re-run, and Jaedyn moved us onto the patched advisory lock. Anne Erdtsieck also contributed a change to let the schema delta decide when an index needs a concurrent build.
JasperFx
JasperFx is just a foundational shared library, but a lot happened there:
The shared vector and hybrid search surface described above
A store-agnostic StubEventStream<T> for unit testing event sourced handlers, with documentation on all three stores
The @jasperfx/event-model-vue renderer moved into the JasperFx repository, next to the Event Model descriptor it draws, and the Event Model now handles services that host more than one model
Hardening the aggregate source generator, including an opt-in build-time assertion that the generator is actually attached
Andre Vieira fixed codegen test failing for every message handler since Wolverine 6.37, and Alan Klimowski fixed a code generation frame ordering issue (and a duplicated service declaration in Wolverine.HTTP)
Fisher
Fisher went from 1.5 to 1.13 in two weeks. Beyond the search work, Fisher now creates its event store tables on first use, supports directory tenancy on Windows, validates tenant ids before turning them into file names, fixes decimal comparisons in LINQ, and requires the source generator with a smoke test of the packed package. Kebin contributed a fix for enlisted sessions with an inline projection registered.
Thank You
The Critter Stack only gets this good because people use it hard (thanks?), tell us when it breaks with actionable error reports, and increasingly send in the fix too. Thank you to everybody who contributed code over the past two weeks:
Ali Yuksekkaya, Anne Erdtsieck, Michael Harris, Travis Kirke, Marko Lahma, Laurence Gillian, Tore Hammervoll, Alan Klimowski, Erik Shafer, Andre Vieira, Jakob Tikjøb Andersen, Joel Reinford, Jaedyn, Mark van der Dam, Raymond Masciarella, Arnel Robles, vpetrevski, Kebin, Marcin Aumiler, Jorge L. Torres M, and Rayan-and-beyond.
And to everyone who filed a good issue with a reproduction, including ArieGato, michielpeeters, raypet-visma, AndreiKopylov, framos-varajo, syserr500, BaharAtNode, Petteroe, zxjon22, r0ss88, bittercoder, BaerMitUmlaut, wpei-infotrack, bohdan-hukivskyi, tychomensing-topicus, and Babu Annamalai: those reports are what drove a lot of this.
No, seriously, the Critter Stack community is as far as we can tell far, far about average for OSS projects in terms of how helpful the community is to help drive and improve the tools.
I write these posts purposely on Fridays (when nobody will read it) just to reflect on where things are at and gather my own thoughts. I will admit that I think there is vastly more feature work in the rear view mirror than there is ahead, but there’s still some things still on the list. But you also never know when some major new software development trends will upend everything you think you think!
CritterWatch 1.1
I’m hoping that CritterWatch 1.1 will be ready for release by the first of October. At this point we’re “officially” feature complete and mostly trying to add more user interface polish and do a lot more manual testing.
The biggest single item is that we’re doing testing in a JasperFx’s client next week who wants to use CritterWatch on an immensely big and complex system (that also uses AWS SQS which is by far and away the whiniest messaging transport). I’m hopeful that getting access to their environment and being able to see CritterWatch admittedly struggle to monitor that massive system will do a great deal for us to get CW 1.1 ready to release.
If everything goes magically well next week, we might be sneaking in a new “Document Explorer” feature set to give you a very comprehensive database viewer for Marten, Polecat, and Fisher built into CritterWatch. That might of course float to 1.2. As it is, CritterWatch 1.1 made our Event Store Explorer feature much more feature rich and flexible to allow you to easily query and view event store data for Marten, Polecat, and Fisher.
You might notice that the project codename “Bobcat” deviates from our Mustelidae naming convention. “Bobcat” has been my working codename for a spiritual and modernized successor to my old Storyteller tool that I’ve been envisioning for longer than the “Critter Stack” naming scheme has existed. I’ve frequently said that my coding superpower is having a much longer attention span than average and Bobcat will maybe be the result of about 17-18 years of work and thinking in this space.
Bobcat is a new MIT-licensed “critter” for improved integration testing and will be the Critter Stack solution to “Spec Driven Development” and a key player in our forthcoming Event Modeling strategy I’ll discuss next.
The big theme is making automated integration testing more successful through:
A “Supervisor” library that uses the Microsoft Testing Platform to “supervise” integration heavy test suites and provide selective and adaptive test retries, flakiness detection, test parallelization, and even manage resources (Docker container resets or even starting an all new process) during tests. I realize that this effort is quite possibly something that’s just useful for Critter Stack development, but it’s already paid off for us!
Using some old ideas (that actually did work even if the project as a whole was a bust!) from Storyteller and FitNesse before that for making data intensive testing easier and provide useful, correlated telemetry for diagnosing test failures
An option for using Gherkin for expressing tests when the Given/When/Then format works or with people who are just comfortable with that approach
Being able to “project” specifications from existing xUnit or TUnit tests through (hopefully) judicious usage of attributes, comments, and source generators
Having your test harness be able to also expose a lot of application instrumentation and telemetry in the outputs for easier AI agent consumption and troubleshooting (again, really an old idea from Storyteller being rebuilt with AI agents in mind).
Some specific add on packages or recipes for HTTP testing with Alba and Spec Driven Development against Event Sourcing with the entire Critter Stack
I’ve thought and discussed a possible markdown authoring approach like ThoughtWork’s Gauge tool, but right now I’m personally leading toward just the combination of Gherkin or the C#/F# tests with projection.
As for the name, I was fascinated as a teen when bobcats started reappearing in our area again (there are actually some very real advantages for wildlife when far fewer people live on farms). But also, as someone with a rural background, people use the name “Bobcat” to refer to any skid steer tool and that carries the connotation of getting stuff done to me:
Event Modeling with the Critter Stack
I’ll share more details soon, but just know that Babu & I are working very hard on a new comprehensive Event Modeling support throughout the Critter Stack. This work and functionality is landing admittedly all over the place with the core JasperFx library defining our official “Event Model” API and that being supported by Wolverine and our event stores including Marten, Polecat, and Fisher. The Bobcat tool up above plays in this too by providing human readable specification rendering and execution. Mostly though, the AI agent coordination, scaffolding, and event model visualization will be from a new tool named “Stoat” that will be included with JasperFx’s AI Skills pack. And of course, our AI Skills tie everything together.
For right now, let me sum that up as:
Generate system code targeting the Critter Stack from a YAML file exported from the Event Modeling AI platform. Shame on me, but I’m not quite sure that’s a standard yet (this?) or just something from their own toolkit
Define an “Event Model” through a Gherkin specification or more likely through .NET code and see a visualization of the event model in a web browser as you work and refine that — and then, of course, generate system code from that model
Generate system code off of the “Event Model” as well as specifications using Bobcat (which could in turn just be riding on xUnit.Net or TUnit)
CritterWatch can visualize an Event Model for your system as it actually is based on the real application based on a combination of the system describing itself and using logging for cause and effect relationships between command messages or HTTP endpoints and the cascading messages or events appended
Our AI Skills and some hard coded scaffolding support in Bobcat make the generated code use the Wolverine and Marten/Polecat/Fisher idioms for low code ceremony and maximize testability
And of course, I’m a big believer in the “Semantic Model” strategy for frameworks, so every possible way you can define the “Event Model” like the import from the EventModelers.AI platform, our Gherkin DSL, a forthcoming C#/F# Fluent Interface we’re going to add to JasperFx.Events, or whatever other 3rd party DSL like ESDB we might support later will be translated first to our very own semantic model. All of our visualization and code generation will target our own semantic model.
That being said, I honestly think that it’s going to be faster in many cases to just write low fidelity C# or F# code in an IDE to generate the event model than it will be to use one of those intermediate modeling DSLs like ESDB or some of the other JSON, XML, or YAML dialects that are popping up. I also think that in the end many folks will probably drive the event model through chatting with an LLM, so we’re investing in our AI Skills to make that as seamless as possible too.
Just to level set, I’m struggling to be too enthusiastic about Event Modeling as a formal method used to specify in detail, then generate systems. I definitely think it can be helpful in a “Uml as Sketch” kind of analytic tool of course and I am a fan of doing interactive Event Storming sessions with domain experts. I’m old enough to remember when Model Driven Development was a complete bust and it’s hard for me to get past the obvious comparison. There are some real downsides to being older and more experienced. It’s natural to compare new ideas to older ideas you’ve already encountered, and sometimes that makes you miss the real differences or value in the new thing. Just something mildly depressing to think about for us older guys still doing software.
Stoat is the codename for a new tool from JasperFx Software that is targeted for AI assisted development, durable agentic memory, visualizing agent activity, coordinating agent activity across code repositories (it’s possible that I’m the only person who cares about that, but I care a lot), and as the glue behind our Event Modeling and Spec Driven Development strategy over all including a user interface for Bobcat specification previews and execution.
There’s a documentation site up, but just know that the website is 100% AI generated with lots of Opus inflected Yoda-speak and many things will change before we inflict Stoat upon the world. I’m hoping that Stoat will have its first release at the same time as CritterWatch 1.1.
We will also be doing a full review of the website to improve the verbiage too of course, but sometimes it’s just easiest now to let an AI tool draft something then mostly rewrite it.
Just know though, that Stoat will be a commercially licensed tool that will be bundled (retroactively too!) with the JasperFx AI Skills pack or CritterWatch.
Wolverine has long had support for message idempotency using Wolverine’s transactional inbox to track messages that have already been handled, but that usage only applies to Wolverine metadata and does nothing from an upstream sender (or let’s be honest, your code) accidentally sending the same logical message.
Recent versions of Wolverine added “logical message deduplication” as well to allow you to specify business logic concerns as the identity for another level of protection.
For example:
The operator clicked Rebuild twice. The console republished the command after a timeout. A scheduling agent pre-published tonight’s 03:00 occurrence yesterday, and the scheduler published it again on the night. Should the projection rebuild four times?
Each of those is a different delivery of the same intent, so each carries a different Envelope.Id and every one of them gets through. What is needed is an identity for the intent, and that is what DeduplicationId is.
We probably should have added this feature ages ago, but we did so in 6.31 to get ready for Wolverine to have a first class (but minimal!) chron message scheduler and use the deduplication id as a way of preventing multiple executions in a chaotic world.We know users are already taking advantage of this because JasperFx has already done some enhancements and optimizations for one of our clients.
You do need to explicitly enable this (because it potentially adds database tables or columns and we do not make mandatory database changes outside of major version releases. Backwards compatibility is a pain, but it’s you know, kind of important).:
Don’t worry, it’s Wolverine, so there are of course some magical policy ways to shorten the configuration.
A logical id is a string so it can be legible to users.
The second message with that id never reaches your handler. It is discarded, acknowledged to the broker, and logged at Information — a duplicate that vanished without a trace would be indistinguishable from a message that was lost.
Where the id comes from
By default a message handler reads Envelope.DeduplicationId. You can point at a member of the message itself instead, so publishers do not have to set DeliveryOptions:
publicstaticclassCreateOrderHandler
{
// Derive the logical id from a member of the message itself rather than
ValueSource.Header reads an envelope header, and ValueSource.Anything uses the chain type’s natural default.
Deriving the id on the publishing side 6.31
Everything above is the receiving half. On the publishing side, asking every call site to remember DeliveryOptions.DeduplicationId is exactly the kind of repetition that eventually gets forgotten at one call site and silently un-protects a message. So a message type can declare its own logical identity once, the same way it can already declare a topic name with [Topic] or a saga id with [SagaIdentity]:
// The message type declares its own logical identity once, and every publisher
Either form is applied as an IEnvelopeRule when the message is routed, so it reaches every transport, the local queues, and the outbox alike. Non-string members are converted with ToString().
Or if you need to, or just want to do so explicitly, you also have this syntax:
If you’ve never heard of it (there’s something like umpteen seasons of a basic cable reality show about it somewhere), just set a timer first to quit before you spend all day reading about it.One of the most kill joy things I have ever read is a somewhat reasonably sounding explanation for the “treasure pit” being a natural sinkhole though.
I was exactly the kind of kid who couldn’t read enough about Stonehenge or the search for the Titanic (which hadn’t been found yet when I was younger) or “Alien Astronauts.” In particular, I have always been fascinated by the Oak Island Mystery where there may (or probably not) be buried pirate treasure on an island off of Nova Scotia. To over simplify it, treasure hunters who have tried to dig up the treasure have continuously stumbled into more problems right after thinking they’ve passed one barrier. Hit a barrier of logs? Rip them up and keep digging, but it turns out there’s another one below it. Dig that one up? The tunnel floods. Send dye down the shaft to try to figure out where and how the water is getting in? There appears to be a tunnel dug to the hole from a cove. Plug that hole up? It appears there’s another tunnel to a different cove! And you get the point.
I have frequently described particularly hard troubleshooting efforts as “Oak Island Problems.” You know, the kind where you pat your back along the way anytime you’ve “moved the exception!” Except in our case, we can generally get to the bottom of it and get things fixed without spending hundreds of years and maybe a dozen lives in accidents trying to find the treasure that may or may not actually exist.
JasperFx Software is closing in on a couple big releases in the next week or two that are worth talking about now to get last minute feedback and maybe additional requests.
Critter Watch 1.1
I’m hoping to have a follow up https://critterwatch.jasperfx.net release that builds on our 1.0 release with a handful of new features:
Cron type scheduled recurring message publishing with a full dashboard for each management as well as matching MCP tools for AI agents. This builds on some recent Wolverine work that added recurring messaging. We looked at Quartz.Net, TickerQ, and Hangfire and decided that we could do what we needed to do for Wolverine and CritterWatch without taking on any explicit coupling. We’ll also be adding guides on what we think the best practices are for integrating Wolverine with the mature job scheduling tools out there, but we won’t be doing formal integration packages.
Stream compacting policies for users to be able to specify recurring background jobs to automatically choose event streams for compacting based on user defined criteria (more than 1,000 events? older than 6 months?). We think this will be a great way to keep Marten/Polecat/Fisher applications performant over time. That will build on top of the Cron messaging.
LLM callouts from CritterWatch alarm detection to use AI tools to potentially decide and carry out amelioration steps. This also builds on recent Wolverine additions for LLM integration.
Scheduling projection rebuilds or subscription rewinds for Marten/Polecat/Fisher so you can have projection rebuilds happen during off hours.
Bringing back the Embedded CritterWatch option to add CritterWatch to a single ASP.Net Core application instead of having to have a separate application. We think this will be very helpful for development time
Moar robustness! One of our earliest users has a phenomenally large system with a complex multi-tenancy strategy that has given CritterWatch internals and user interface quite a workout
Formal support for Event Modeling diagrams of configured and observed system behavior. To be more clear, this is Critter Watch being able to show you an Event Modeling visualization of the system as it actually is according to Wolverine configuration and observing the cause and effect between command messages and events appended or other messages being cascaded. This is part of our larger effort toward Event Modeling support.
Event Modeling Support and Spec Driven Development
Alright, this one is an arc across the entire Critter Stack to enable people who want to use Event Modeling, then use AI tools to scaffold or build applications using those models. We’re also working on a first class Spec Driven Development story for Critter Stack applications with and in addition to the Event Modeling tooling.
Here’s a diagram of the three main ways we’re looking to support Event Modeling and Spec Driven Development across the Critter Stack, with a pair of newer critter tools that aren’t quite to 1.0 yet in “Stoat” and “Bobcat” (more on these below).
The major pieces are:
A “Semantic Model” in our low level JasperFx.Events library that models everything there is to a vertical slice according to Event Modeling semantics
A new library called “Bobcat” that’s going to be our main Event Modeling visualization tool and Gherkin specification executor. It’s also a spiritual successor to my much older Storyteller project and has a lot of the same DNA for hopefully making automated integration testing more successful in enterprise systems. Bobcat also has command line helpers for generating skeleton Wolverine code for the event slice model
Our JasperFx curated AI Skills that help constrain your AI tools to generate the cleanest and most idiomatic possible code you can with the Critter Stack as well as helping you choose the best Wolverine or Marten options for whatever your system needs. The AI Skills will also help you utilize the Critter Stack test automation support.
A new tool named “Stoat” (no website yet, but we’re working on it!) that’s our “everything that makes AI usage more successful” tool including a feature set inspired by KurrentDb’s Capacitor that will add durable memory across your AI agents and coordination and visibility across AI agents. Stoat will be the controller for our entire Spec Driven Development story. Our thinking right now is that Stoat will be a commercial tool that will be bundled into AI Skills or CritterWatch purchases from JasperFx Software.
Imagine a couple approaches to modeling an event driven system:
Use the EventModelers.AI tooling for modeling. Stoat will be able to take the YAML file of the model exported from EventModelers.AI, then invoke AI agents and Bobcat to scaffold and develop a Critter Stack system from the model
Using nothing but Gherkin to define the slice model for your system, including BDD style specifications, and have Bobcat create a running Event Modeling visualization of your system as well as enabling Stoat to again execute a development plan based on that model
Lastly, and here’s where the Critter Stack is going to diverge quite a bit from seemingly the rest of the Event Sourcing community, a model where you can just write at least a skeleton of C# or F# types for events, read models, and command messages, then use a lightweight fluent interface in code to generate the Event Modeling visualization for easy review. I personally do not believe that the attempts to create intermediate DSLs for modeling event sourced systems like ESDM are going to be successful and that the Critter Stack’s very low code ceremony actually makes it easier to just write C# or F# code with a real IDE instead of futzing with DSLs.
But, to that last point, I think we’re going to be well set up to potentially add support later for all those DSLs and YAML formats and even that really gnarly 2000’s era looking XML format going around in conferences by translating those to our own event slice model and working from there.
And just to be clear, Stoat & the AI Skills and therefore our full Spec Driven Development strategy including our support for the EventModelers platform will be commercial add ons to the CritterStack. We’re still committed to “Open Core,” but these items are going to fall out of the MIT licensed core.
I’m of the opinion that everybody’s approach to AI assisted development is exactly what their prior opinions were about software development. I was much more heavily influenced by Extreme Programming back in the day and have always been very dubious about any kind of “Model Driven Development” and that has carried through to being dubious about a lot of these new event modeling or spec driven development approaches. That’s also why you’re going to see us emphasize BDD style specifications more than modeling approaches.
Extreme Scalability for Marten and Wolverine
JasperFx has an ongoing effort with a large client that involves some rather extreme scalability requirements (10’s or 100’s of billions of events in a single system, 500+ different PostgreSQL databases, about 250 unique message types and that many HTTP web services. And a modular monolith to boot). We’ve already done quite a bit, with more ideas still to come. JasperFx and our client will be publishing a white paper sometime this year laying out all the challenges they faced and what we did to be able to enable their scale.
I’m pretty excited for this to all come to fruition.
Random Things?
More AI integration into Wolverine with the possibility of it maybe being a more durable option in place of Microsoft’s Agent Framework?
DuckDb integration for Marten, with Polecat and Fisher coming later
We shipped Critter Stack AI Skills 1.10.0 today. Eleven new skills, which brings the catalog (so far) to 102 across all the Critter Stack tools and trying to cover every use case and common troubleshooting need we can think of.
If you haven’t run across these yet, the short version is that AI Skills are structured documentation written for the coding agent rather than for you. Not API reference — the agent can already read that — but the accumulated “here’s what this actually means, here’s the trap, here’s what to check next” that otherwise only exists in the heads of the people who built the thing. I wrote up a whole session of an agent using them against a real codebase a couple of days ago if you want to see them working.
Our AI Skills will also get your AI agents to use Wolverine, Marten, Alba, and all the other tools in the most idiomatic way possible that should lead to more terse code and in many cases, more performant code. They’ll also help your AI agents write more testable code and the automated tests that go with that using all the Critter Stack test utilities we’ve built over the years. And lastly, the AI Skills help your agents to understand test failures and to utilize all the observability capabilities of the Critter Stack with some help from CritterWatch’s MCP support as well as all the command line diagnostics that come along with the Critter Stack.
In short, we’re throwing every possible bit of AI spaghetti up against the wall and the AI Skills are kind of the glue of all our AI related tools and approaches.
Here’s what’s new.
Sagas, from both directions
Wolverine Sagas was a big hole in the AI Skills before this release, and now we’ve got AI Skills coverage on the valid usages of Saga handler signatures and for using Sagas from HTTP endpoints covers the part that trips up almost everybody who tries it. In Wolverine.HTTP, the first return value of an endpoint method is the response body. Which means if you return a Saga, congratulations — you’ve just serialized your saga state to the caller. The Saga has to be a later tuple member, and there’s an [EmptyResponse] shape for when you don’t want a body at all. That one fact reshapes the whole topic, and it’s exactly the kind of thing that is obvious in retrospect and expensive in the moment.
Troubleshooting Projections
There are two new projection troubleshooting skills today.
Troubleshooting a projection is the general one: your projection is behind, stalled, or producing a document that’s just plain wrong, and you need to figure out which. The CritterWatch one is the fleet version of running applications in production — diagnose_projection to rank shards worst-first, then run_projection_stepper to replay your actual production projection code over a slice of real events and read the before/after state at every single step.
For development time, we added a new command line tool called projection-run to JasperFx.Events and shipped it in Marten and Polecat. Then we added a skill that knows how to use the new “projection stepper” CLI tool and tested it by pointing an agent at a projection I’d deliberately broken and watching where it got stuck.
Marten: versioning, archiving, and search
We filled in some important gaps for Marten usage in the AI Skills this time around:
Event versioning and upcasting — renaming an event type, moving a namespace, adding a field, and the three flavors of upcaster for when the change isn’t additive.
Archiving and stream compaction has the fact I most want people to internalize: archiving alone doesn’t shrink anything.ArchiveStream sets a flag. It’s UseArchivedStreamPartitioning that moves archived events onto separate physical storage and actually buys you the query performance — and it quietly weakens a stream-identity guarantee on the way, which you should know about before you turn it on.
gRPC with Wolverine, covering the service-shim-to-bus-to-handler flow, streaming and its cancellation contract, and RPC deduplication.
MCP servers for your own app — twenty shipped tools across Marten.Mcp, Polecat.Mcp and WolverineFx.Mcp that expose your application to an agent: query event streams, fetch aggregate state, read daemon status and routing diagnostics, scaffold a vertical slice.
And CritterWatch alerts — how to acknowledge, snooze, or clear an alert with the right audit attribution, and how metrics alerts are actually decided underneath.
Getting them
The AI Skills are available standalone or bundled with CritterWatch:
Like probably all software tool companies, JasperFx is working very hard to create a compelling story about the utilization of AI-assisted development with our tools. The MCP support — and command line tools too! — in CritterWatch is a major part of our AI strategy.
So far JasperFx Software has mostly been showing off CritterWatch as a user interface tool that you’ll use to peruse and explore what’s happening in your system.
Today, let’s shift to how CritterWatch empowers your AI tools to understand and even administer your running suite of Critter Stack applications through its MCP tools. I should note a few more things before we jump into the sample usages:
Everything exposed through the MCP tools in CritterWatch is also available through a command line package
The MCP tools are gated by your CritterWatch license
We have invested, and will continue to invest, in making JasperFx’s curated AI Skills know exactly how to take advantage of both the MCP and CLI tools — as well as all the other command line diagnostics already built into the Critter Stack!
Before we get into the real details, just know that everything below is a real recorded session from earlier today: Claude, pointed at the CritterWatch MCP endpoint on my dev machine, driving our full Aspire-orchestrated sample fleet that we use to develop and test CritterWatch itself — 25 monitored services, PostgreSQL and SQL Server event stores, RabbitMQ, Azure Service Bus emulator, and AWS SQS via LocalStack, all chattering away. The trouble in this session was manufactured on purpose with CritterWatch’s built-in chaos monkey tools (which are themselves MCP tools — more on that later), because a demo that waits around for production to genuinely catch fire makes for a long blog post. But every tool call, every JSON response, and every “wait, that’s not what I expected” moment is the real thing, lightly trimmed for length — and the blockquoted Agent replies you’ll see are how Claude actually wrote the raw JSON back up for me, tables and recommendations included. That’s not me grading my own homework — it’s the same MCP surface you’d point your own agent at.
Two minutes of setup
The MCP server rides along in the CritterWatch console host. If you’re already running CritterWatch, you mount it with two lines:
csharp
builder.Services.AddCritterWatchMcp();
// ...
app.MapCritterWatchMcp(); // mounts at /api/mcp
That registers 48 tools — 21 read tools and 27 action tools — over MCP’s streamable HTTP transport, deliberately configured stateless so every tool invocation sees the actual caller’s identity for authorization (a subtle thing that matters a lot once RBAC is in play — see the end of this post). If your host is already composing an MCP server from the per-tool packages, there’s a chaining overload that folds CritterWatch’s tools onto the same endpoint alongside them.
Pointing an agent at it is a config stanza, not a project. For Claude Desktop or Claude Code:
jsonc
{
"mcpServers": {
"critterwatch": {
"url": "http://localhost:5173/api/mcp",
"transport": "streamableHttp"
}
}
}
There’s a quick start for the consumer side that covers MCP Inspector, Claude Desktop, and rolling your own client. That’s the whole setup story. Now the fun part.
“What’s going on in production?”
I would like to point out that CritterWatch’s AI abilities should also be valuable at development time, and we’re investing quite a bit in our AI-assisted development story and expect CritterWatch to be a big part of that.
The session opened the way every on-call shift opens — with the vaguest possible question.
And here’s the part that raw JSON dashboards never give you — the agent turned that into an answer, not a data dump:
Agent: Your fleet looks healthy right now:
Monitored services
25
Wolverine nodes
29 (4 distinct hosts)
Endpoints
1,049 — all Active or Accepting, none stopped, no senders latched
Broker connections
14 / 14 healthy
Nothing looks off at the fleet level. Want me to drill into a specific service, check for active alerts, or look at projection lag anywhere?
Twenty-five services, over a thousand endpoints, fourteen healthy broker connections — read, interpreted, and handed back with the next three sensible questions already offered. Nothing was wrong — yet. So I gave it something to find.
Dead letter triage, the whole loop
The dead letter queue is where I think the MCP story goes from “neat” to “oh, I’d actually use this.” CritterWatch has a first-class Dead Letters page, but the MCP surface closes the full loop for an agent: two read tools to discover and explain what died, and two action tools to do something about it — without a human ever pasting envelope ids around.
For this session I used our internal chaos monkey tool to write 40 realistic dead letters into TripService’s message store — varied message types, varied exception types, and a configurable fraction seeded as genuinely recoverable, so “replay what will succeed” is a real decision rather than a scripted gesture. Then:
"... grouped by message type × exception type, largest first ..."
],
"databasesAnnounced": 1,
"databasesAnswered": 1,
"partial": false
}
This is the triage view: counts grouped by message type and exception type, per message database. Here’s how the agent presented it back to me — and notice that it doesn’t just reformat, it diagnoses:
Agent: TripService is carrying 40 dead letters, all in the trips PostgreSQL store. Every announced database answered (1 of 1), so these counts are the complete picture, not a floor.
The shape here matters: many exception types spread across many message types reads like transient infrastructure trouble (timeouts, connection resets, contention) rather than a single poison message. Most of these should succeed on replay. Want me to pull the individual envelopes and try?
That “the shape here matters” reasoning isn’t something I prompted for — an agent (or a human) reads the pattern: lots of exception types across lots of message types smells like a dependency being down; one message type with one exception smells like a poison message. And do notice databasesAnnounced versus databasesAnswered — the reads fan out across every physical message database the service owns and tell you explicitly whether everyone reported. An empty result where only three of eight stores answered is not an empty queue, and the tool refuses to let an agent confuse the two. That distinction exists because of a real production failure mode where a console rendered “no dead letters found” over a queue quietly holding 42 of them.
From there the agent pulled individual envelopes — with ids, exception messages, and source URIs. Here’s one of them, verbatim:
The replay is fanned out to the service’s own message store and re-handled there — CritterWatch relays the command over the same Wolverine messaging backbone it uses for everything else. A minute later the agent re-summarized:
And then reported back with exactly the partition you’d want from a colleague:
Agent: Replay complete. Here’s how it shook out:
25 of 40 drained — every TimeoutException, DataException, ArithmeticException, and ExternalException cleared on retry, which confirms those were transient infrastructure failures.
15 bounced straight back to the queue, and every survivor is a System.InvalidOperationException — the handler rejects these messages deterministically. Replaying them again will produce the same result. These look like genuine poison: either the payload is malformed or the handler has a bug for this input shape.
I’d recommend discarding the 15 (or holding them for a developer to inspect first). There is no undo on a discard — say the word.
Genuinely poisonous messages don’t get better because you asked nicely, so I said the word:
Agent: Done — 15 discarded. Final check: summarize_dead_letters reads 0 dead letters with all databases reporting ("partial": false), so that’s a true zero, not a zero with a shrug.
Forty dead letters to zero, with the agent narrating what it replayed, what it didn’t and why, asking before the irreversible step, and verifying its own cleanup. The ids flow from the read tools straight into the action tools; no console tab was harmed in the making of this triage.
“What happened with message 08df07a4-…?”
As part of the effort toward building CritterWatch, we purposely improved Wolverine’s OpenTelemetry output to enable the troubleshooting we’re showing in this section.
This one’s my favorite, because it’s the question you actually ask in an incident review. Somebody hands you a message id from a log line and wants to know: where did this thing come from, what did it do, and what happened because of it?
Wolverine has emitted rich OpenTelemetry spans for years — every send, receive, and handler execution, tagged with messaging.message_id, messaging.conversation_id, messaging.message_type, the destination URI, and the handler type. CritterWatch’s trace tools sit on top of whatever tracing backend you already run (Jaeger in my dev fleet; DataDog and Application Insights providers ship in the box) through one provider abstraction, so the agent never needs to know or care which backend answers.
Me: What happened with message 08df07a4-3ded-0488-baa4-8b734d770000?
Look at that root span, because it’s telling you something genuinely interesting: this message wasn’t born in a controller or a handler. It was born inside a Marten async daemon page execution — the projection daemon chewing through events #2030–2034 on the trips database. The agent then pulled the full span tree with get_trace — spans come back flat with ParentSpanId edges, and every Wolverine span carries the semantic tags:
"... 7 more spans: the daemon page, sibling sends, sibling handler executions ..."
] }
And here’s the answer the agent actually gave — the raw spans reconstructed into the story a human wants:
Agent: Here’s the full life of message 08df07a4-3ded-0488:
Born — inside a Marten async projection daemon page on postgresql://localhost/postgres/trips, while it processed events #2030–2034 (a 15.3ms page). This message is a consequence of event processing, not a user action.
Sent — TripService dispatched it as a TripMessages.ContinueTrip command over RabbitMQ (49-byte payload).
Handled — TripPublisher.ContinueTripHandler executed it on TripPublisher with status OK.
Siblings — the same parent operation emitted three more ContinueTrip messages in the same batch (…0474, …0514, …0100); all three were also handled successfully.
So: five events were processed, four commands came out of it, all four landed cleanly across two services. Nothing about this message needs your attention.
Events appended, the message they caused, the handler that ran, and the sibling messages that rode along — the entire causal neighborhood of one message id, reconstructed from spans your system was already emitting. And when I asked the structural version of the same question — “where does ContinueTrip fit in the overall workflow?” — the agent reached for describe_lifecycle, which returns the type’s complete lifecycle stitched across every monitored service: publisher → transport → handler → cascaded messages → appended events → projections, rendered as both structured JSON and a ready-to-paste Mermaid sequence diagram. The observed instance and the structural map, side by side.
Chasing a poisoned projection
At least in the early days of Marten, this was frequently one of the hardest things to troubleshoot, but CritterWatch is going to help a great deal.
Now for the event sourcing crowd. Async projections fail in a uniquely annoying way: some event, somewhere in a stream, makes your Apply method throw, and the failure surfaces a long way from the cause. So I armed CritterWatch’s deterministic projection poison against TripService — set_chaos_monkey_projection_poison(eventTypeName: "TripStarted") — which makes every apply of that event type throw. (Why deterministic instead of the random failure-rate knob? Because with a rate, the alert, the dead-letter drill-in, and the projection stepper each land on a different random event, which is exactly what a diagnosis story can’t use. The poison makes every surface name the same event.)
Within a couple of minutes the fleet got loud, and the agent’s list_active_alerts("TripService") came back looking like a proper incident:
"exceptionMessage": "Failure to apply event #5784 Id(95dca7ce-5330-4f1d-9f18-151f91670d59)" }
Which the agent rolled up into an incident assessment I’d happily paste into a ticket as-is:
Agent: This is a poison-event pattern, not an outage. Here’s the evidence chain:
All 2,100 dead letters are one message type (TripMaintenanceCheck) failing with one exception type (ApplyEventException) — the opposite shape from this morning’s transient batch.
The exception is thrown by the event store’s apply pipeline, not by handler logic — each of these handlers folds the Trip aggregate from its event stream, and the fold itself is what’s dying.
The failure names the exact event: event #5784 on stream 95dca7ce-5330-4f1d-9f18-151f91670d59, a TripStarted apply.
Every incoming maintenance check is going to keep dead-lettering until the apply is fixed — replaying now would just bounce all 2,100 straight back. Recommended next step: open this stream in the projection stepper and watch the fold up to event #5784 to see exactly what state + event data the Apply method chokes on.
Read that middle bullet again: the agent distinguished “your handler threw” from “your aggregate fold threw” purely from the exception’s origin — which is exactly the distinction that decides where you go looking for the bug. And this is precisely the moment the projection stepper exists for: give it the projection and the stream, and it replays the events one at a time in the console, showing you the projected document’s state before and after each event — so you watch the state fold step by step right up until the apply that blows up, with the failing event’s data sitting in front of you. No more “add a Console.WriteLine to the Apply method and rebuild.”
The remediation ran back through MCP, in order: clear the poison (in real life: ship the fix to your Apply method), then page through and replay all 2,100 maintenance checks — and the agent narrated its own progress as it went, so I could watch the drain without asking:
Agent: Poison cleared. Replaying the backlog in pages of 200 and confirming each drains before continuing:
replayed 200 — ~1,900 remaining
replayed 200 — ~1,700 remaining
replayed 200 — ~1,500 remaining
...
replayed 200 — ~300 remaining
replayed 100 — queue empty ✓
All 2,100 replayed and handled successfully — the fix holds under the full backlog, not just a sample.
Then restart_projection("TripService", "Distance:All") to bounce the projection agent, whose restart command CritterWatch routes to whichever node currently owns that agent:
Which the agent handed back as a briefing rather than a JSON blob:
Agent:11 active alerts across 6 of your 25 services — 4 critical, 7 warnings:
Service
Critical
Warning
PolecatTripService
2
—
HelpDesk
1
3
TripService
1
1
MTTripService
—
1
Trip3Service
—
1
IncidentService
—
1
The TripService critical is the dead-letter rate alert from this morning’s drill — that one’s explained and should decay on its own. PolecatTripService’s two criticals are the ones I’d look at next if you want to keep going.
That’s a real fleet with real background noise, not a sanitized screenshot — eleven active alerts across six services, most of them the ordinary grumbling of a dev fleet that’s been abused all afternoon. list_active_alerts filters by service or severity, get_alert drills into one, and the action side — acknowledge_alert, snooze_alert, clear_alert — flows through the same event-sourced alert streams the console UI writes to. When the agent ran acknowledge_alert("alert:TripService:MetricsDlqRate:*") on the dead-letter-rate alert (my fault, see above), that acknowledgment showed up in the console’s alert timeline like any operator action, attributed and auditable.
One alert in that list turned out to be my favorite kind of detail: a warning that TripService’s scheduled-job poller was sitting on a growing backlog of scheduled envelopes. That wasn’t chaos I’d injected — that was CritterWatch correctly flagging genuine congestion in the sample fleet’s trip lifecycle, which schedules a TripMaintenanceCheck fifteen seconds after every trip starts. The demo rig got caught by its own monitoring. I’ll take it.
Read models, tenants, and performance while we’re here
A few quick hits from the same session, because the read surface goes well past the headline acts:
get_projection_lag("TripService") returned all seven projection shards with their sequence positions, assigned nodes, agent status, pause reasons, and any captured exceptions — the projection-monitoring view as data.
query_documents("TripService", "trip") paged through the 6,356 live Trip documents in the service’s Marten store — the document store explorer as an MCP tool, which also happily queries the daemon’s DeadLetterEvent documents when you’re hunting skipped events.
On the multi-tenant services, list_document_types enumerated each tenant-sharded store’s types, and the tenant action tools — add_tenant, enable_tenant, disable_tenant, and friends — drive the same runtime tenant lifecycle the console UI does, fanned out across every registered tenant source.
list_backlog_hotspots and get_backlog_state roll up incoming/outgoing/scheduled/dead-letter depths per store, fleet-wide, for “is anything drowning?” at a glance.
That document query deserves one concrete look, because “the agent can read your read models” is easy to skim past. Asked for a sample of live trips, it came back with:
Agent: 6,356 Trip documents in the store. Something worth a second look, though: the trips I sampled all show "Traveled": 0 despite being active for a while — if that field is supposed to accumulate, the projection feeding it may not be receiving the travel events at all.
That last observation is the kind of thing you get for free when the entity reading your data can also reason about it — and yes, it was right: the sample fleet’s travel pipeline really was stalled behind that scheduled-message backlog the alerting had flagged. The agent connected two findings I hadn’t connected myself.
The other half: AI Skills
Here’s the thing I want to be really clear about, because I think it’s the actual strategy and not just a feature list. The MCP tools expose the live, queryable system. But a pile of tools doesn’t make an agent good at operations — an agent also needs to know the discipline. That’s what the JasperFx AI Skills are for, and the two are built as complementary halves of one story.
The skills are curated, versioned instruction sets your coding agent loads on demand — there are dozens covering Wolverine, Marten, Polecat, testing, and messaging patterns, and a set specifically for operating CritterWatch: dead-letter triage, lifecycle diagnostics, routing diagnostics, service actions, and setup. The DLQ triage skill, for instance, doesn’t just list the four tools — it teaches the loop (summarize → query → act, in that order, with the ids flowing through), teaches an agent to read grouped counts as symptoms (“many exception types across many message types reads as a dependency being down; one message type with one exception reads as a poison message”), and drills in the non-negotiable rule I mentioned earlier: an empty result with partial: true means some stores did not report — never answer “the queue is empty.” Every time the agent in this session checked databasesAnswered before declaring victory, that was the skill talking.
We hold ourselves to a standing rule internally: any time CritterWatch exposes new information through an MCP tool, the paired skill work ships with it. A tool with no skill coverage is an under-leveraged tool.
What else can the agent do?
The session above touched maybe half the catalog. The rest of the toolbox, quickly:
Event stores and read models — the document explorer tools (list_document_types, query_documents, get_document) across every monitored service’s store, plus the Event Store Explorer and projection stepper in the console for stream-level spelunking.
Alerts — list, summarize, drill in, acknowledge, snooze, clear; alert thresholds themselves are configurable in the console.
Scheduled jobs — the console’s scheduled messages view tracks every service’s scheduled envelope backlog (and as you saw above, the alerting watches the poller’s drain rate for you).
Projection monitoring and operations — get_projection_lag for the shard-by-shard truth, then pause_projection / restart_projection / rebuild_projection / eject_projection, all tenant-scopable.
Performance — backlog state and hotspots fleet-wide, projection lag, plus the OpenTelemetry trace tools riding your existing Jaeger / DataDog / App Insights backend.
Message routing — explain_message_routing answers “where would this message go and why?” from the actual runtime routing rules, which beats reasoning about conventions from memory every single time.
Listeners and services — pause, restart, and drain listeners; evict a defunct service from the console.
Chaos engineering — the same chaos monkey tools I used to rig this demo: failure rates, slow handlers, deterministic projection poison, and dead-letter seeding for game days against staging.
One important word about the gates
Everything in this post is a paid-tier capability — every MCP tool checks the CritterWatch license before doing anything, and an unlicensed host answers every tool call with a polite LicenseMissing envelope rather than your production data. And beyond licensing, the action tools run through CritterWatch’s opt-in RBAC: every mutating tool is gated on a named capability (dlq.replay, dlq.discard, chaos-monkey.configure, and so on) scoped to the target service as a resource — so “the agent may replay dead letters on TripService but touch nothing on the billing service” is a policy you can actually write. Even the dead-letter reads carry their own capability, because dead letters contain message bodies, and “may look at business data” deserves a separate grant from “may act on it.” RBAC ships off by default for the low-ceremony getting-started path, and turns on when you’re ready to hand an agent real credentials in a real environment. The stateless MCP transport exists precisely so those checks always see the current caller, not whoever happened to open the session.
Handing an AI agent a control surface for production is the kind of thing that should make you a little nervous. It makes me a little nervous, and we built it — which is exactly why the license gate, the capability model, and the per-service scoping aren’t bolted on after the fact.
Epilogue: the alerts that wouldn’t clear
Dogfooding is good, and this issue is already fixed for our upcoming 1.1 release.
One loose thread, because I promised you the “wait, that’s not what I expected” moments too. Sharp-eyed readers may have noticed something off in the poisoned-projection section: those AgentDown alerts claimed the projection agents hadn’t sent a heartbeat in over four minutes — when the poison had only been armed for two. And after the remediation, they lingered longer than they should have, over a daemon that a direct SQL query showed advancing every single second.
The symptoms were real. The mechanism was not what anyone in the room — human or agent — guessed at first. Chasing it down with these same tools against CritterWatch’s own telemetry pipeline turned into a genuinely fun piece of distributed-systems forensics, complete with a Postgres deadlock storm, a docker daemon picking the worst possible moment to restart, and a feedback loop that made a merely-slow pipeline read as a dead one. We caught it entirely in the course of putting this post and its samples together, and the fix is already in for the upcoming CritterWatch 1.1 release. The full diagnosis — and what it taught us about building telemetry lanes that shed load instead of aging it — is the next post.
Summary
CritterWatch’s MCP server is two lines of host code and a config stanza in your agent, and it exposes 48 tools across dead letters, alerts, health, performance, traces, routing, documents, projections, tenants, listeners, and chaos.
The dead-letter loop — summarize, query, replay, discard — closes end to end without a human ferrying envelope ids around.
“What happened with message X” is answerable from the OpenTelemetry spans Wolverine already emits, through whatever tracing backend you already run.
When a projection sours, the agent finds the exact poisoned event from the dead-letter evidence, and the projection stepper shows you the state folding right up to the failure.
The AI Skills teach your agent the operational discipline the tools alone can’t; the two ship as halves of one story, and the skills are bundled with CritterWatch Professional and Enterprise.
If you want to try this against your own Wolverine or Marten system:
I read a post from Paul Stack last week I thought was interesting titled AI Broke the Assumptions Behind CI. I was maybe much more influenced by Extreme Programming (XP) back in the day, so I’d disagree a little bit with his characterization of CI being something tied to the pull request workflow in GitHub et al, but let’s talk more about this.
The original point of CI was really just to be constantly getting feedback on your code by building and running tests against it as you change code and adjust as needed if CI found problems. With the radically different part of XP at the time being that you actually wrote tests!
The general idea behind XP was to be able to work faster and more adaptively by providing your team with effective feedback loops to guide and correct the adaptation as you worked. I’d say after all these years that CI is still very important — but it’s time for a serious rethink in our new AI powered world order.
Back to what Paul was getting at in his post, here’s the CI process I admittedly use for Marten, Wolverine, or other Critter Stack tool development:
Write code locally and be constantly running the tests most closely related to whatever I’m changing
Create pull requests — which we do in no small part just for traceability more so that as a code review tool (GitHub release notes generation is tied to pull requests and that’s just very helpful)
Lazily allow GitHub Actions execute all kinds of test suites and smoke test harnesses against the pull request while I go off and worry about something else
Merge pull requests when CI goes green or fix broken tests
And that’s mostly worked out until just recently. However, with the extreme load that’s come from all of us yahoos using AI agents to code so much faster, GitHub Actions are very noticeably slower or flat out unreliable on the worst days. And of course, to make things worse, we’ve added a lot more tests and CI actions than we ever tried to do before.
Granted, the Critter Stack work I do is especially problematic in this regard because so much of what we do involves slow running tests that execute against databases and message brokers. We’re also having to constantly build up and tear down .NET applications in memory. Especially for us, but maybe for you too since so many of us depend on now overloaded CI servers, there’s a very real problem that depending on CI builds running on a remote box has become a sometimes unacceptable bottleneck.
To balance a “good enough” safety net and for getting things done, I’ll mirror Paul Stack’s post and say that in some cases I’m:
Doing trunk based development like it’s 2007 and Subversion is the latest hotness!
Strictly using local testing with a lighter weight set of suites for commits, and trying to be very selective of what subset of tests are executing based on the changes in flight.
Using the full blown “HeavyGate” of test suites to do pushes to the remote main
In public projects where there’s value in using pull requests just for tracking, I think we’re going to have to get more creative about test run filtering to avoid running test suites that aren’t relevant to the changes in flight to get pull requests in without hours of delays.
On a side note, with CI builds being so slow, that’s forced us to be much more aggressive about stomping out flaky or otherwise unreliable tests so CI is far more consistent for us. I’ve also added a new critter named “Bobcat” (the project is actually old enough that I started it before we adopted the current Critter Stack naming scheme) that among other things, uses the Microsoft Testing Platform to supervise test runs and selectively do test retries, process restarts, and even hard Docker resets based on known test flakes. That’s been hugely helpful both locally where heavy development can break down when Docker containers have run too long in tests and also for CI where retrying a CI failure just in case it’s just a test flake is just too damn slow now.
Anyway, AI assisted development continues to take up more of my gray matter than I wished it did and I don’t think the adaptation is going to change any time soon.
It’s apparently time to play another game of .NET developers feeling consternation because of an OSS project’s attempt to be more sustainable. JasperFx Software, the company I founded around the “Critter Stack” tools is committed to an “Open Core” model. What that means for us is that:
The main libraries and tools like Marten and Wolverine will remain MIT licensed, i.e. free and open
JasperFx also has its commercial AI Skills offerings for the Critter Stack and now our commercial CritterWatch tool. With another set of commercial tools related to AI assisted development and Event Modeling coming soon (those will be part of the same license as CritterWatch)
At this point I think that JasperFx Software has already established that we have a viable business model and that we’re going to be able to sustain our “Open Core” model going forward.
We still get occasional friction from potential clients and users that we’ll do the same OSS “rug pull” that some other projects did by adopting a commercial license on their newer versions even though I’ve said “Open Core” in public a half a thousand times. I do not particularly enjoy that level of cynicism about OSS that frequently crops up in the ,NET community.
What I would say to folks out there is that tools like Marten and Wolverine are just not viable as side projects. A huge amount of our functionality in Marten for scalability only existed after I founded JasperFx and started working directly with clients every day. Wolverine’s leadership election which undergirds a lot of our advanced blue/green deployment support and scalability would not have been possible to build and constantly curate without me being full time on the Critter Stack. I would ask our users to have some awareness for how much time it takes to evolve and maintain their OSS tools.
Yes, the existence of AI tools makes it tempting to think you can just vibe code replacements for your 3rd party dependencies over a rainy weekend, but you have to also understand how much hardening widely used OSS tools get from being beaten up by users and having to adapt to a world of technical irregularities like database outages, Rabbit MQ quietly dropping connections, database overloading, network hiccups, database administrators unexpectedly sending a kill signal to a PostgreSQL database that turns out to create gaps in sequences (and wasn’t that one fun), and not to mention all the crazy edge cases we’ve had to face from Kubernetes doing Kubernetes things.
To sum this all up, you can’t just vibe code replacements for quite a bit of this, these kinds of tools achieve deep quality through a lot of usage, feedback, and adaptation over time — and all of that takes a lot of time and a long attention span.
Hell, I’ve personally had to make several improvements to code subsystems in Marten and Wolverine in the last month that I thought were “done” and as stable as they could possibly be because new users in new circumstances proved otherwise
And just because I might get asked about this, my friend Ian Cooper wrote about this too, but maybe coming from a different perspective as an OSS maintainer. I partially agree with some of that and I’ll respectfully disagree with other parts and just leave it at that.
I’m obviously sympathetic to the Polly maintainers, and based on this exchange with one of the creators of the OSMF, my initial inclination is to pay the OSMF fee from JasperFx Software because of our commercialization of Marten, Polecat, and Fisherthrough support plans (those projects are still MIT licensed folks!) and the transitive dependency that CritterWatch has on Polly through those other libraries:
Like I said, I’m sympathetic to the Polly maintainers, and we’ll try to be above board with them here. But, if there’s even the slightest bit of hesitation from our current or potential customers about the Polly license, we’ll replace our relatively small usage of Polly with something new in our foundational JasperFx library and remove Polly entirely. I’m not enthusiastic about doing that because the Polly.Core dependency is in our public API and pulling that out would require us to either do a major version release or cheat on SemVer rules — which we really hate to do without very good reason.
I after all have a fiduciary responsibility to my “shareholders” to make JasperFx a sustainable financial success.
Wolverine has its own resiliency features, so Polly isn’t a concern there at least. Our document database and event store applications do use Polly for resiliency against transient errors though, and that’s what would need to change. I’d guess that most of our users don’t even realize that’s there, so maybe the switchover won’t be that big a deal if we decide to go that way.