diff options
| author | blasty <blasty@local> | 2026-08-07 23:16:30 +0200 |
|---|---|---|
| committer | blasty <blasty@local> | 2026-08-07 23:16:30 +0200 |
| commit | 41b3710d5b6f9be24430fd55c46c83f9ae3b8d83 (patch) | |
| tree | ed3517997a3248310c904c3a57ea053257162ac2 /idatui/drive.py | |
| parent | Export findings as markdown (Ctrl+E), and the journal that makes it true (diff) | |
| download | ida-tui-41b3710d5b6f9be24430fd55c46c83f9ae3b8d83.tar.gz ida-tui-41b3710d5b6f9be24430fd55c46c83f9ae3b8d83.tar.xz ida-tui-41b3710d5b6f9be24430fd55c46c83f9ae3b8d83.zip | |
Ctrl+F: search the whole database, by text or by bytes
`/` only ever searched the lines of the view you were in. This adds the
search you actually need on a binary: over the entire database, either
through the rendered disassembly or through the image.
* **text** matches the line as displayed, whitespace-normalised, so
`call cs:` finds `call cs:getenv_ptr` (IDA's column padding is not
something anyone types). Smartcase; `regex` available over RPC.
* **bytes** is IDA's own `find_bytes`, so the pattern language people
already know works unchanged: hex pairs, `?` wildcards for a whole byte
or one nibble (`48 8? ?? 24`), quoted literals (`"Hello", 0`). Commas,
no separators (`488B05C3`) and ragged spacing all normalise.
**Which mode you meant is guessed, and the guess is biased on purpose.**
`dead`, `add`, `cafe` and `ff` are valid hex AND ordinary things to search
for, so a bare hex-looking word stays TEXT; nobody types `48 8b ?? c3`
meaning prose. `hex:`/`text:` prefixes and F2 override it.
The subtle case is a *typo* in a byte pattern. `48 zz c3` first fell
through to a text search and reported "no match" — indistinguishable from
"those bytes are not in this binary", which is the most misleading answer
a search can give. Now any query whose tokens are all byte-sized is
treated as bytes, and a bad token is refused BY NAME. IDA does the same
thing quietly (find_bytes answers a malformed pattern with zero hits and
no error), so the validation lives in Program.search, not just in the UI.
Enter searches, then Enter opens the highlighted hit; the title says which
it will do, because a database-wide scan is far too slow to run on every
keystroke like the other palettes. Navigation goes to the item head — a
byte match can start mid-instruction — and the status names the exact
address.
Also: the `find` RPC verb and `drive find`, which is the one an agent
wants (`drive find '48 8b ?? c3'`).
idatui/search.py holds the classification and is pure, so the whole
question of "what did they mean" is tested offline: tests/test_search.py,
35 checks, 0.1s. Pilot scenario db_search covers the UI end to end.
Full suite: 890 passed, 0 failed, 51.3s.
Diffstat (limited to 'idatui/drive.py')
| -rw-r--r-- | idatui/drive.py | 16 |
1 files changed, 16 insertions, 0 deletions
diff --git a/idatui/drive.py b/idatui/drive.py index 6ba13c6..6e5c21a 100644 --- a/idatui/drive.py +++ b/idatui/drive.py @@ -296,6 +296,21 @@ def cmd_save(c, args): return " saved" +def cmd_find(c, args): + """find <query...> -- search the database; bytes if it looks like bytes.""" + if not args: + raise SystemExit("usage: find <text | 48 8b ?? c3 | hex:...>") + r = c.call("find", query=" ".join(args)) + hits = r.get("hits", []) + out = [f" [{r.get('mode')}] {len(hits)}{'+' if r.get('truncated') else ''} hits"] + for h in hits[:40]: + out.append(f" {h['addr']} {(h.get('func') or h.get('seg') or ''):<20.20} " + f"{h.get('line', '')}") + if len(hits) > 40: + out.append(f" … {len(hits) - 40} more") + return "\n".join(out) + + def cmd_export(c, args): """export [path] -- write the session's findings as markdown.""" r = c.call("export", **({"path": args[0]} if args else {})) @@ -324,6 +339,7 @@ COMMANDS = { "rename": cmd_rename, "mv": cmd_mv, "note": cmd_note, "retype": cmd_retype, "save": cmd_save, "screen": cmd_screen, "raw": cmd_raw, "define": cmd_define, "syms": cmd_syms, "fmt": cmd_fmt, "export": cmd_export, + "find": cmd_find, "binaries": cmd_binaries, "switch": cmd_switch, } |
