--- name: capemon-developer description: Expert capability for navigating, modifying, and extending the capemon malware monitoring codebase. Includes deep knowledge of Windows API hooking, PE structures, and the CAPEv2 sandbox architecture. --- # Capemon Skills `capemon` is a monitoring and instrumentation engine designed for malware analysis, configuration extraction, and payload recovery. It acts as the core injection component for the CAPEv2 sandbox. ## Core Capabilities ### 1. API Hooking & Monitoring `capemon` implements an extensive hooking engine derived from `cuckoomon-modified`, providing deep visibility into application behavior across multiple subsystems: - **Process & Thread Management:** Monitoring creation, termination, and manipulation of processes and threads. - **File System Operations:** Tracking file creation, deletion, reading, and writing. - **Registry Activity:** Capturing configuration changes and persistence mechanisms. - **Network Communication:** Intercepting socket operations, DNS queries, and high-level protocol activity (HTTP, etc.). - **Cryptography:** Extracting keys and monitoring encryption/decryption routines. - **Synchronization & Services:** Monitoring mutexes, events, and Windows Service interactions. - **Windows Management Instrumentation (WMI):** Intercepting WMI queries used for anti-analysis or reconnaissance. - **Scripting Engines:** Specific hooks for VBScript and other language runtimes. ### 2. Debugging & Tracing `capemon` implements an in-process debugger independent of Windows debugging interfaces, but harnessing the capabilities of the processor: - **Hardware breakpoints:** Four breakpoints bp0-bp3 that can be set on execute, read or write - **Software breakpoints:** Unlimited INT3 or 'CC' breakpoints overwriting instruction byte - **Single-step:** Tracing allows instruction-level capture enhanced with configurable step-over, trace-length, register changes, function names, strings & more - **Actions:** Configurable actions allow control flow manipulation with skipped or taken jumps, arbitrary register changes or jumps, string capture, dumps, scans & more - **Programmable:** Debugger configurable either on submission with simple text options or via dynamic YARA signature scans during unpacking or detonation - **Integration:** Hooking engine integrated with optional behavior log output & breakpoints set on return from hooked APIs (break-on-return) - **Stealth:** Debugger does not rely upon Windows interface and thus evades detection by a slew of interface-related indicators, with additional stealth from hook-based protections ### 3. Automated Unpacking 'capemon' implements an unpacking engine using a combination of techniques - **Memory region tracking:** Regions of memory revealed through indicators of execution, allocation or protection are tracked - **Early capture:** Multiple possible triggers allow payload capture at earliest moment often resulting in working unpacked samples - **Injection capture:** Strong coverage of injection techniques for inter-process payload capture - **PE unmapping:** Integrated Scylla engine allows capture of memory or file-mapped PE images in memory - **Shellcode dumping:** Shellcode & non-PE regions equally captured as payloads - **Import Reconstruction:** Repairing Import Address Tables (IAT) to create functional dumped executables. - **AMSI Dumping:** Intercepting and dumping buffers passed to the Antimalware Scan Interface (AMSI). ### 4. Config Extraction Automated Static & Dynamic malware configuration extraction relies on 'capemon' capabilities - **Static extraction:** Typically reliant upon capemon's unpacking or process dump capture before static parsing - **Dynamic extraction:** When parser implementation is onerous, dynamic capture of decrypted configs can be performed by debugger via YARA signature ### 5. YARA integration Integration of YARA for in-memory scanning - **Dynamic configuration:** Sandbox configuration such as hooking exclusions or options implemented during detonation - **Debugger programming:** Precise dynamic breakpoint address resolution using YARA signatures & cape-specific metadata - **Unpacking engine integration:** Dynamic scanning of all memory regions prior to unpacking capture - **Function resolution:** Allows dynamic address resolution for APIs or functions for hooking or other purposes ## Technical Foundations - **Platform:** Windows (x86 and x64). - **Hooking Method:** Inline hooking of Win32 and Native APIs (NTAPI). - **Debugger:** Native in-process 'self' debugging utilising minimal OS interfaces & hardware capabilities (breakpoint, single-step) - **Dependencies:** - `distorm` for instruction decoding. - `libyara` for pattern matching. - `Scylla` for PE reconstruction. - `bson` for data serialization. ## Architectural Idioms (MANDATORY - read before designing any feature) capemon has one established way of doing each of the following. New code MUST plug into these paths; parallel mechanisms will be rejected in review. ### A. Detection = YARA, not C scanners - Byte-pattern, magic-value or header detection is a YARA rule. Monitor-wide detections go in the built-in `InternalYara[]` string (`CAPE/YaraHarness.c`); family/config detections go in the analyzer yara directory. - Do NOT write `for (p = start; p < end; p++) if (*(DWORD*)p == MAGIC)` loops. Express header validation in the rule condition (`uint32(@s[i] + N)`, `for any i in (1..#s) : (...)`). ### B. Where detection runs = the unpacking engine's region scans - `CAPE_post_init()` runs BEFORE unpacking. Anything packed (UPX, crypters, shellcode loaders) is not yet visible there. - `YaraScan(Address, Size)` is already invoked on the initial image (`CAPE.c`), on tracked regions when they are executed or change protection (`ProcessTrackedRegion` etc. in `CAPE.c`), and on debugger-driven scans (`Trace.c`). A rule added to the rule set fires on all of these automatically. - Never tie a feature to the `ImageBase` global or `GetModuleHandle(NULL)` alone. ### C. Action dispatch = `cape_options` metadata - A rule triggers behaviour via `meta: cape_options = "..."`. `YaraCallback` parses it with `user_data` = base of the scanned region, and `ParseOptionLine` resolves `$string` references to match addresses. - A new action is a new keyword handled in `YaraCallback`, which calls the feature with an explicit address (region base and/or `base + Match->offset`). - Features take the base address as a parameter: no globals, and no re-discovery of what YARA has already located. - Regions are re-scanned; expect repeated hits on the same region and make features idempotent (dedup by address). ### D. Debugger ownership - Each breakpoint owns its handler: `SetBreakpoint(..., Callback)` for hardware breakpoints, `SetSoftwareBreakpoint(BPs, Address, Callback)` for INT3. - NEVER modify `SoftwareBreakpointHandler`, `SoftwareBreakpointCallback`, `SingleStepHandler` or `CAPEExceptionFilter` to dispatch a feature. Doing so hijacks every other debugger consumer (traces, YARA `bp` options, syscall breakpoints). - Single-step state is per thread; do not overwrite the global `SingleStepHandler` to service a feature. ### E. Behaviour log: one `LOQ_*` call per API call or breakpoint hit - Every hook or breakpoint emits at most ONE `LOQ_*` record per hit. Gather every relevant field first (function name, arguments, buffers, resolved names), then emit them together in a single call with a multi-field format string (e.g. `"sSS"`). - Do NOT log a generic "function called" record and then one record per parameter. Unreadable arguments are logged as empty or zero-length values inside the same record. - A call that spans entry and return (e.g. output buffers) logs once, at the point where the data is available (usually on return). Stay silent at entry. - Bulk metadata (module info, file lists, dependency trees) is one record per object, not one per item. Filter out noise (stdlib, dependencies already listed elsewhere) and cap the size. - `DebugOutput` goes to the debug log, not the behaviour log, but it should not emit per-item floods either. ### F. Design review checklist (answer before writing code) 1. Does an existing mechanism already do this? Check `YaraScan` callers, `YaraCallback` options, `SetBreakpoint*`, tracked regions, `DumpRegion`. 2. Does it work on a UPX-packed sample? On a non-PE (shellcode) region? 3. Are addresses passed as arguments rather than read from globals? 4. Does it modify a shared handler? If yes, redesign. 5. Is it gated by a config option and documented in `docs/configuration.md`? 6. Does each hit produce at most one behaviour log record? ### Case study: Go hooking (kevoreilly review, PR #181) - **Rejected:** `GoRecoverSymbols(GetModuleHandle(NULL))` called from `CAPE_post_init()`, a hand-written C pclntab scanner, and `GoBreakpointHandler` dispatched from `SoftwareBreakpointHandler` for every software breakpoint. Missed all UPX-packed Go samples. - **Accepted:** `rule golang` in `InternalYara` -> `cape_options = "golang"` -> `YaraCallback` -> `GoRecoverSymbols(base, ...)`; Go hooks registered with `SetSoftwareBreakpoint(..., GoBreakpointHandler)`. ## Engineering & Documentation Mandates - **Always update `@docs/configuration.md`:** Whenever a new configurable option is introduced to the engine (such as `log-format`, `sleep-skip-seconds`, etc.), you must immediately append its documentation details to the appropriate table inside the configuration reference document to ensure the user and the system documentation are fully up-to-date. - **Always Fetch and Merge Upstream (`upstream/capemon`):** Before starting any development task, creating a new branch, or preparing changes, you MUST always fetch from upstream (`git fetch upstream`) and merge `upstream/capemon` into your working branch so all work builds upon the latest commits. Never work on stale code. If any merge conflicts arise, you MUST resolve all conflicts completely and verify that compilation and functionality remain intact across all targets (Win32, x64). ## Source Control Workflow ### Upstream Synchronization & Conflict Resolution Mandate > [!IMPORTANT] > **Always Work on Latest Upstream (`upstream/capemon`):** > 1. **Always Fetch Upstream**: Before creating a new branch, starting any task, or making changes, always ensure your working branch is updated with the latest upstream commits: > ```bash > git fetch upstream > git merge upstream/capemon > ``` > Ensure the local base branch and working branches are strictly in sync with `upstream/capemon` (`kevoreilly/capemon`). > 2. **Resolve All Conflicts**: If any conflicts arise when merging upstream changes into an active branch or worktree, inspect each conflicting file, resolve all conflicts thoroughly, and verify that the resulting code compiles cleanly for both Win32 and x64. Never leave conflict markers or unresolved states. > 3. **Sync Before PR & Finalization**: Before pushing commits or finalizing PR branches, fetch and merge `upstream/capemon` again to guarantee clean, fast-forwardable or conflict-free integration. > 4. **Maintainer Commits First on Shared PRs**: Maintainers (e.g. kevoreilly) push directly to PR head branches. Before applying review fixes, fetch the PR head (`agent_worktree.py update ` or `git pull --ff-only`) and confirm the maintainer's latest commit is in your branch. Never force-push over a PR head; never start fixes on a head that predates the maintainer's push. ### Isolated Checkouts for Review and Testing `capemon` uses the fork convention: `origin` is your own fork, `upstream` is `kevoreilly/capemon`, and the default branch is `capemon` (not `master`). Reviewing someone's PR or testing a branch therefore means juggling two remotes, and doing it with `git checkout` in your working clone risks the uncommitted work most clones carry. Use `.gemini/skills/capemon-developer/scripts/agent_worktree.py`, a wrapper around `git worktree` that handles the remote and PR plumbing. It is Python standard library only - no virtualenv, no dependencies, and it does not care that this is a C project. ```bash python .gemini/skills/capemon-developer/scripts/agent_worktree.py new --pr 123 python .gemini/skills/capemon-developer/scripts/agent_worktree.py new --branch some-topic-branch python .gemini/skills/capemon-developer/scripts/agent_worktree.py new --from upstream/capemon --name scratch python .gemini/skills/capemon-developer/scripts/agent_worktree.py list python .gemini/skills/capemon-developer/scripts/agent_worktree.py path pr123 python .gemini/skills/capemon-developer/scripts/agent_worktree.py update pr123 python .gemini/skills/capemon-developer/scripts/agent_worktree.py remove pr123 python .gemini/skills/capemon-developer/scripts/agent_worktree.py cleanup python .gemini/skills/capemon-developer/scripts/agent_worktree.py info ``` Behaviour relevant to this fork layout: * The canonical repository is read from `upstream` when it exists, so `new --pr ` queries `kevoreilly/capemon` rather than your fork. * `new --pr` asks `gh` which fork the head branch lives in and fetches from the matching remote, or straight from the fork URL when no remote matches. * `new --branch` tries `origin`, then `upstream`, then any other remote, and reports which one supplied the branch. Force one with `--remote`. * Worktrees default to `~/.cache/agent-worktrees/capemon/`; override with `--base-dir`, `--path` or `$AGENT_WORKTREE_DIR`. * Only worktrees the tool created are ever removed - they are tagged inside the repository's git admin directory, so `git status` stays clean and a hand-made `git worktree add` is never touched by `cleanup`. * `remove` and `cleanup` refuse to discard uncommitted changes or unpushed commits, and the main worktree can never be removed. `--force` overrides. Add `--json` to any command for scripted use. `gh` is required only for `--pr`. > The same script is maintained in the CAPEv2 repository as > `utils/agent_worktree.py`; keep the two copies in sync when changing it. ### Building a Worktree A worktree is a full checkout, so the MSBuild commands in the next section work unchanged inside one - point the solution path at the worktree instead of your main clone. Build output stays in the worktree and disappears with it. ## Build & Compilation Guide ### 1. Locating MSBuild On a standard Windows development machine, MSBuild may not be present in the global `PATH`. You can locate it using PowerShell by running a query over the standard Microsoft Visual Studio or Build Tools installation directories: ```powershell Get-ChildItem -Path "C:\Program Files", "C:\Program Files (x86)" -Filter "MSBuild.exe" -Recurse -ErrorAction SilentlyContinue | Select-Object -ExpandProperty FullName ``` Typical installation paths include: * **Visual Studio 2022 Build Tools (32-bit/64-bit host):** `C:\Program Files (x86)\Microsoft Visual Studio\2022\BuildTools\MSBuild\Current\Bin\MSBuild.exe` * **Visual Studio 2022 Community Edition:** `C:\Program Files (x86)\Microsoft Visual Studio\2022\Community\MSBuild\Current\Bin\MSBuild.exe` ### 2. Compilation Targets and Toolset Overrides The `capemon` solution specifies the legacy Visual Studio 2017 (`v141`) platform toolset. If your local build system only has Visual Studio 2022 (`v143`) installed, you can compile successfully by dynamically overriding the platform toolset and disabling Whole Program Optimization (`LTCG` / Link-Time Code Generation) to prevent linker mismatches against precompiled static `.lib` dependencies (like `libyara`). #### Compiling Win32 (x86) Release Target: ```powershell $msbuild = "C:\Program Files (x86)\Microsoft Visual Studio\2022\BuildTools\MSBuild\Current\Bin\MSBuild.exe" & $msbuild /m /p:Configuration=Release /p:Platform=Win32 /p:PlatformToolset=v143 /p:WholeProgramOptimization=false capemon.sln ``` #### Compiling x64 (64-bit) Release Target: ```powershell $msbuild = "C:\Program Files (x86)\Microsoft Visual Studio\2022\BuildTools\MSBuild\Current\Bin\MSBuild.exe" & $msbuild /m /p:Configuration=Release /p:Platform=x64 /p:PlatformToolset=v143 /p:WholeProgramOptimization=false capemon.sln ``` ### 3. C++ Compilation & Include Order Guidelines When developing or integrating C++ components (such as the `.NET` profiler) into the `capemon` C codebase, adhere to these guidelines to prevent compiler/linker errors: * **Preventing Winsock Redefinition Conflicts**: Always include `WinSock2.h` before `windows.h` inside C++ files or headers to prevent legacy definitions from being pulled in by default: ```cpp #ifdef _MSC_VER #include #endif #include ``` * **Required Include Order for .NET Profiler Headers**: `corprof.h` relies on definitions from `cor.h` and `corhdr.h`. To avoid compilation/syntax errors, use this exact order: ```cpp #include #include #include #include ``` Additionally, add `#pragma comment(lib, "corguids.lib")` in your source files to link the standard GUID definitions for COM callbacks and profiler interfaces. * **C++ Keyword and Redefinition Conflicts (`hooks.h`)**: Never include `hooks.h` inside C++ files. `hooks.h` contains parameter declarations using `this` (which is a C++ keyword) and tentative global variable declarations (which cause `LNK2005` duplicate symbol errors in C++). If you need to access monitor/dump functions like `SetCapeMetaData` and `DumpMemoryRaw`, declare them manually as `extern "C"` rather than including `hooks.h` or `CAPE/CAPE.h`. * **C++ Type-Safety for Allocations (`alloc.h`)**: Since C++ does not support implicit conversion from `void*`, any allocation calls from `alloc.h` inline functions (e.g., `cm_alloc`, `cm_calloc`, `cm_strdup`) inside C++ compilation contexts must be explicitly cast to `(char*)` or the appropriate pointer type.