aboutsummaryrefslogtreecommitdiffstats
diff options
context:
space:
mode:
authorblasty <blasty@local>2026-08-09 22:16:09 +0200
committerblasty <blasty@local>2026-08-09 22:16:09 +0200
commit57bbc2e2959b56dc24c09c9a3f8b09a357c60c9e (patch)
tree10c00069012d72450484f17e63c11c3068b90fc8
parentDrop the settrace workaround: ida-codemode 0.3.2 deleted the hook (diff)
downloadida-tui-57bbc2e2959b56dc24c09c9a3f8b09a357c60c9e.tar.gz
ida-tui-57bbc2e2959b56dc24c09c9a3f8b09a357c60c9e.tar.xz
ida-tui-57bbc2e2959b56dc24c09c9a3f8b09a357c60c9e.zip
Docs: mark the two upstream perf findings fixed in 0.3.2
CODEMODE_UPSTREAM items 1 (timeout_trace) and 2 (to_jsonable) both landed upstream; note it at the top, on each item and in the priority table, and flag that items 3-9 are not re-verified against 0.3.2. SPEED.md described the settrace workaround and its IDATUI_CODEMODE_TRACE knob as current -- both are gone, and its backend table predates 0.3.2.
-rw-r--r--.fastfeedback/SPEED.md23
-rw-r--r--docs/CODEMODE_UPSTREAM.md27
2 files changed, 41 insertions, 9 deletions
diff --git a/.fastfeedback/SPEED.md b/.fastfeedback/SPEED.md
index 996aefe..4274145 100644
--- a/.fastfeedback/SPEED.md
+++ b/.fastfeedback/SPEED.md
@@ -214,12 +214,23 @@ IDA boot — 7 processes, each opening its own database.
| `decomp_map` | 45.8ms | 47.9ms | 1.0x |
| empty round trip | ~0.07ms | **2.0ms** | the floor |
-Two things dominated and are fixed (see `idatui/codemode_client.py:_script`):
-`to_jsonable` walking every returned object (snippets now return one
-pre-serialised JSON string), and `sys.settrace` — the runtime installs a trace
-that returns itself, i.e. LINE tracing in every frame, which made
-`ida_bytes.get_flags` 52x slower than native. The snippet detaches it and
-restores it in a finally; `IDATUI_CODEMODE_TRACE=1` keeps the stock behaviour.
+Two things dominated, and **both are now fixed UPSTREAM in ida-codemode 0.3.2**,
+so neither is our problem any more:
+
+- `sys.settrace` — the runtime used to install a trace that returned itself, i.e.
+ LINE tracing in every frame, making `ida_bytes.get_flags` 52x slower than
+ native. 0.3.2 deletes the hook (deadline is now a C-level thread interrupt).
+- `to_jsonable` walking every returned object. 0.3.2's `serialization.dumps_json`
+ fast-paths straight to the C encoder.
+
+Our two client-side workarounds were re-measured against 0.3.2 and buy **0.99x**
+and **0.97x** respectively — nothing. The settrace strip (and its
+`IDATUI_CODEMODE_TRACE` escape hatch) is **deleted**; `_PACK_EPILOGUE` is kept
+only to pin encoder settings. Re-measure with
+`PYTHONPATH=. ~/ida-venv/bin/python experiments/bench_pack_trace.py`.
+
+**Numbers in the table above predate 0.3.2** and were taken with both workarounds
+active; treat them as historical.
**What is left is the 2ms round-trip floor, and it is NOT ours.** Measured against
the same worker, same connection:
diff --git a/docs/CODEMODE_UPSTREAM.md b/docs/CODEMODE_UPSTREAM.md
index f903599..f8d79e0 100644
--- a/docs/CODEMODE_UPSTREAM.md
+++ b/docs/CODEMODE_UPSTREAM.md
@@ -11,6 +11,16 @@ unnecessary.
**Environment:** ida-codemode 0.3.1, IDA 9.4 (idalib), Linux, single managed
worker backend, quiet box. Target for timings: `targets/echo` unless stated.
+> **Status against 0.3.2 (upstream `93e8aad`).** Items **1 and 2 are FIXED** —
+> the two that mattered most. `sys.settrace` is gone from the runtime entirely
+> (the deadline is now a C-level thread interrupt, `runtime._interrupt_thread`),
+> and `serialization.dumps_json` hands results straight to the C encoder,
+> falling back to the `to_jsonable` walker only for values `json.dumps` rejects.
+> Both of our workarounds re-measured at **0.99x and 0.97x** on 0.3.2 — i.e.
+> nothing. The settrace strip has been deleted; `_PACK_EPILOGUE` is kept only
+> for encoder determinism. Harness: `experiments/bench_pack_trace.py`.
+> Items 3–9 are **not** re-verified against 0.3.2 and may still stand.
+
**What the client does**, for scale: it renders a continuous disassembly listing,
pseudocode, a CFG graph view and a hex view, paging over the database as the user
scrolls. It is latency-sensitive in a way an agent-driven MCP client is not — a
@@ -20,7 +30,9 @@ keypress must repaint. It issues ~1–8 operations per user action.
## 1. `timeout_trace` enables line tracing in every frame — 52x on IDA calls
-**Highest-impact item by a wide margin.**
+**Highest-impact item by a wide margin.** — ✅ **FIXED in 0.3.2.** The runtime no
+longer installs a trace hook at all; cancellation is a C-level thread interrupt.
+Our `sys.settrace(None)` workaround is deleted as of `a5137fe`.
`runtime.py` wraps every `execute_python` in `sys.settrace(timeout_trace)` to
enforce the deadline. `timeout_trace` ends with `return timeout_trace`, and
@@ -75,6 +87,11 @@ and do the same, which is an argument for fixing it in the runtime.
## 2. `to_jsonable` dominates any large result
+✅ **FIXED in 0.3.2**, via the first suggested fix below: `serialization.dumps_json`
+calls `json.dumps(value, default=to_jsonable)`, so a JSON-safe result never enters
+the Python walker. Our packing workaround now measures 0.97x and is retained only
+to pin encoder settings, not for speed.
+
`execute_python` runs `to_jsonable()` over whatever the snippet returns. Our
answers are already JSON-safe and they are big — a 200-row listing page is
roughly 10k small objects.
@@ -250,8 +267,8 @@ added a test asserting our kwargs are a subset of
| # | item | impact | fixable by you? |
|---|---|---|---|
-| 1 | `timeout_trace` line tracing | 52x on IDA calls, 10x on real operations | yes, one line |
-| 2 | `to_jsonable` on large results | 114x on serialisation | yes |
+| ~~1~~ | ~~`timeout_trace` line tracing~~ | ~~52x on IDA calls~~ | ✅ fixed in 0.3.2 |
+| ~~2~~ | ~~`to_jsonable` on large results~~ | ~~114x on serialisation~~ | ✅ fixed in 0.3.2 |
| 7 | no change/revision counter | correctness for shared editing | yes, cheap |
| 4 | loader switches fatal on reopen | crashes, hard to diagnose | yes |
| 5 | replaced/deleted IDB under lease | silent hang | yes |
@@ -265,5 +282,9 @@ the private worker it replaced" and "the port is within 2x, and faster on severa
operations". Both are in the runtime, not in client code — which is why they are
worth fixing centrally rather than leaving each client to rediscover.
+**Both landed in 0.3.2**, and both client-side workarounds could then be measured
+at parity and retired. That is the outcome this document was written for; the
+remaining items (3–9) have not been re-checked against 0.3.2.
+
Happy to supply the benchmark harness (it is backend-agnostic and runs against
both our old worker and Code Mode), or to test a patch.