# 🚀 New Release: mcp-windbg 1.5.0

> Failed opens no longer leak debugger sessions, a timed-out command on a dump gets cancelled instead of wedging the session, and completion markers are unique per session. Plus what I learned measuring six community pull requests against a real debugger.

- Published: 2026-10-09
- Author: Sven Scharmentke
- Canonical: https://svnscha.de/posts/mcp-windbg-1-5-0-release/
- Tags: ai, debugging, mcp-windbg, crash-analysis, kernel-debugging, community, release

---

**mcp-windbg** 1.5.0 is out. It's a bug fix release that came almost entirely from a backlog of community reports, and working through it taught me something about reviewing contributions that I want to write down while it's fresh.

Two people did the digging: [@xiaozhu1337](https://github.com/xiaozhu1337) filed two detailed issues and four pull requests after an overnight debugging session went wrong, and [@robster7674](https://github.com/robster7674) found that a slow command could wedge a dump session. Every fix in this release traces back to one of them.

I did not merge all of those pull requests, and the reason is the interesting part, so it gets its own section below.

## Failed Opens Left Debugger Sessions Behind

All four `open_*` tools registered their session before running the initial analysis. If that analysis then failed, say a `!peb` that timed out, you got an error with no session ID, while the session stayed in the server's registry with its debugger still running. You had no ID, so you could not close it. Retrying added another one.

[@xiaozhu1337](https://github.com/xiaozhu1337) reported this in [#125](https://github.com/svnscha/mcp-windbg/issues/125) with a reproduction that does not need a debugger at all, and it worked first try:

```text
attempt 1: Command timed out after 10 seconds: .lastevent
attempt 2: Command timed out after 10 seconds: .lastevent

sessions left registered : 2
shutdown() calls         : 0
```

An open now owns its debugger until initialization, the initial analysis, and output filtering have all succeeded. On a failure, or if you cancel the request, that call's session is rolled back and shut down, and other sessions are left alone. The design is theirs; I changed one thing about it.

A rollback does not resume a live target. Closing a session normally lets the target run again, which is what `close_kd_session`'s `resume` parameter is for. But an open that failed never handed you a session ID, so resuming on its behalf is a side effect you cannot opt out of. Since CTRL+B on a user-mode remote detaches *and* resumes, a rollback now sends nothing at all and leaves the server's target exactly as it was.

## A Timed-Out Command Wedged a Dump Session

A command that outran its timeout was only cancelled on a live target. The reasoning was that a dump has nothing to break into, which sounds right and is wrong: the debugger is still executing the command. On a dump it ran to completion while every command you sent afterwards queued behind it and timed out too. The session was stuck until you closed it.

That is easy to reach now that `open_kd_dump` exists, because `!process 0 7` over a large kernel dump runs for minutes. [@robster7674](https://github.com/robster7674) hit it on real dumps and sent [#124](https://github.com/svnscha/mcp-windbg/pull/124).

Here is the same sequence on a real dump with a three second timeout, before and after. The slow command times out either way; the question is what happens to the next one:

```text
before:  the command times out, then .lastevent ALSO times out after 20s
after:   the command is cancelled in 3.00s, .lastevent returns in 0.00s
```

CTRL+BREAK does cancel a command the engine is running on a static dump. The engine reports `User interrupted operation` and goes back to the prompt.

One detail I changed. If the cancel cannot get the session back to a known prompt, the two session kinds now behave differently, because your options differ. A live target keeps its session, since `send_ctrl_break` can still rescue it. A dump session is closed, because `send_ctrl_break` refuses a dump, so the old advice to "break in manually" was something you could not act on.

## Markers That Could Finish Another Client's Command

mcp-windbg knows a command has finished by echoing a marker after it. Those markers were numbered from a counter that restarts at 1 in every session, so every session used the same names. Attach several clients to one shared `-remote` debug server, where everyone sees everyone's output, and one client's marker could complete another client's command. `..._1` also matched inside `..._10`.

Each session now carries a random nonce, and the match is anchored at the end of the line.

## Smaller Fixes

- The output kept for one command is bounded now, by line count and per line, with a notice saying how it was cut. The existing limits only ever bounded the diagnostic snapshot in a timeout message, not the output actually stored, so a runaway command could grow without limit.
- `.logopen /u` quotes its path. More on this below, because it was worse than it looked.
- A `close_*` that cannot shut its debugger down reports the failure instead of success, so a debugger still holding your target is visible rather than silent.
- `--verbose` writes to stderr instead of stdout, which on the stdio transport carries the MCP protocol. It also installs a logging handler at all, which it never did, so the server's own log messages actually appear now.

## What I Measured Before Merging

Here is the part I found genuinely instructive. Of the six pull requests, four had a premise that did not survive contact with a real debugger. All four were written confidently, all four passed their own tests, and the tests were the problem.

**The completion marker is never alone on its line.** One pull request made marker matching strict: the line had to be exactly the marker. That looks like an obvious tightening. But `cdb` writes its prompt without a trailing newline, so the prompt and the marker share a line:

```text
'0:000> COMMAND_COMPLETED_MARKER_1\n'      # cdb on a user-mode dump
'8: kd> COMMAND_COMPLETED_MARKER_1\n'      # kd on a kernel dump
```

Strict matching rejects every real marker. Measured against the dumps committed in the repo, the released code opens a dump in 0.2 seconds and that branch timed out after 15 seconds, on both `cdb` and `kd`. It passed 107 tests because the test fake emitted a bare marker, which no real debugger does, and one test actively asserted that a prompt-prefixed marker must *not* count. The suite certified the bug.

The awkward part is that the loose substring match it replaced was load-bearing. It is the only reason a prompt-prefixed line completes a command at all. The defect and the working mechanism were the same expression.

So the fix that shipped keeps the marker unique and anchors the match, and never looks at the prompt. Parsing prompt forms works until you meet one you did not plan for: `0:000>`, `1:001:x86>`, `8: kd>` with a varying processor number, `lkd>`, `*BUSY* 0:000>`, and `?:???>` when a remote target has no current process yet. The test fake now emits a prompt-prefixed marker and can be given any of those forms.

**`-y` does not override `_NT_SYMBOL_PATH`.** Another pull request fixed what it described as the dump directory shadowing your configured symbol server. The debugger appends the environment variable after `-y` instead:

```text
_NT_SYMBOL_PATH=SRV*C:\cache*https://msdl.microsoft.com/download/symbols
cdb -z <dump> -y C:\dumpdir -c ".sympath;q"

Symbol search path is: C:\dumpdir;SRV*C:\cache*https://msdl.microsoft.com/download/symbols
```

Both entries, in the order you would want. The change was a no-op, and it would have broken an existing test on any machine that has `_NT_SYMBOL_PATH` set, which is most machines that do Windows debugging.

**A quoting bug nobody was looking for.** While checking a consolidated branch I found something real that was not in any report. `.logopen /u` was called without quotes. Given a path with a space in it, `cdb` does not simply fail:

```text
.logopen /u C:\temp with spaces\out.log
  -> Opened log file 'C:\temp'
     ^ Extra character error in '.logopen /u C:\temp with spaces\out.log'
```

It opens a log at the truncated prefix, leaves that file behind, and never creates the one you asked for. Since that log is how mcp-windbg reads output on a Chinese or Japanese code page, the fix for [#102](https://github.com/svnscha/mcp-windbg/issues/102) was silently off for anyone whose temp path contains a space, such as a user name with a space in it. My own machine resolves `TEMP` to its 8.3 short form, so neither my testing nor CI could see it.

I also introduced a bug of my own while fixing the honest-close behaviour: a failed rollback started replacing the error that explained why the open failed. Three tests failed to catch it, including the contributor's, because `str()` on an MCP error embeds the chained traceback, so searching the whole string finds the original error even when it has been replaced. Asserting on the first line catches it.

The lesson I took is narrow and practical: for anything that talks to a real debugger, a green hermetic suite is not evidence. The fakes encode what we believe the debugger does, and that belief is what needs checking.

## Tested Against a Real Kernel Target and a Chinese Code Page

CI cannot cover a multibyte code page, a live kernel target, or a remote server whose target is running, and those are exactly where these bugs hid. So before releasing I ran the candidate on a Hyper-V VM:

- On code page 936 with a Chinese system locale, the full suite gave the same 207 passed as my Western machine, the Unicode log engaged, and `.echo` with Chinese characters round-tripped intact.
- With `TEMP` set to `C:\temp with spaces`, the log opened at that path and no truncated stray file was left behind. That is the quoting fix proven on the thing it is about.
- Both kernel scenarios passed against the live target over KDNET.

One process change worth mentioning: CI never ran on any of these six pull requests. They came from first-time contributors, so every workflow run sat waiting for approval, and none of the branches had ever been tested by the project's own matrix. That is now configured so contributor pull requests get the full matrix automatically. It would have caught the strict-marker regression on its own, since CI opens a real kernel minidump.

## Still Open

The original report also found something I have not fixed. When you connect to a debug server, its target may legitimately already be running, and the handshake mcp-windbg uses is not proof that the target is ready to answer. Two attempts at fixing it each regressed a different case, and the honest reason is that there is no automated coverage of a running remote target at all: every live test opens a stopped one. Adding that test comes first, so this is tracked in [#140](https://github.com/svnscha/mcp-windbg/issues/140) for the next release rather than rushed into this one.

## Upgrading

```bash
pip install -U mcp-windbg
```

If you use the Claude Code plugin:

```text
/plugin marketplace update mcp-windbg
/plugin update mcp-windbg-uvx@mcp-windbg
```

The full details are in the [v1.5.0 release notes](https://github.com/svnscha/mcp-windbg/releases/tag/v1.5.0). Thanks again to [@xiaozhu1337](https://github.com/xiaozhu1337) and [@robster7674](https://github.com/robster7674): closing a pull request is not the same as ignoring it, and this release would not exist without their reports. If something behaves differently than you expect, [open an issue](https://github.com/svnscha/mcp-windbg/issues).
