--- name: ti-edgeai-import-model description: > Get a trained neural network that exports to a static-shape ONNX (a classifier, a segmentation model, or a detector with a TIDL meta-architecture such as YOLOX/SSD/YOLOv5-style) running on a TI Edge AI board's C7x/MMA accelerator (TDA4VM, AM68A, AM69A, AM67A, AM62A): preflight the ONNX, fold input normalisation, compile and int8-calibrate with edgeai-tidl-tools in Docker (tools tag matched to the board SDK), verify offload and accuracy in host emulation, check board-vs-host outputs, package model/ artifacts/ param.yaml dataset.yaml, deploy to /opt/model_zoo and smoke-test with edgeai-gst-apps. Use this skill whenever the user wants to deploy, convert, compile, quantize, port or "put my model on" a TI Jacinto/AM6xA board, mentions TIDL compilation, artifacts, tidl_net.bin, meta_arch_type, a prototxt, calibration images, int8 accuracy loss on the TI device, or asks why a compiled model fails on the board - for any architecture within those contracts, and even if they only say "deploy to the TI device". license: MIT metadata: version: 0.1.0 tested_on: "TDA4VM, Edge AI SDK 11.0: a YOLOX-tiny TI-lite detector and a float-input classifier through the full flow; other SoCs untested" --- # Import a Model into TI Edge AI When this skill is active, **read the reference named in each phase before running it.** **Scope**: a model whose ONNX export has static input shapes and mostly supported operators, in one of three task contracts (below). Out of scope: dynamic shapes (export fixed ones), operators TIDL cannot run (they fall back to ARM or block offload), models with several inputs or non-image inputs (not tested), and TFLite/TVM flows. The three task types share one flow and differ only in the compile options and `param.yaml`: | Task | Compile extras | Output the board app expects | Run end to end | |---|---|---|---| | detection with a TIDL meta-architecture (YOLOX/YOLOv5/7 type 6, SSD 3, YOLOv3 4, RetinaNet 5, YOLOv8 8) | `--meta --meta-arch-type N` | `dets [N,5]` + `labels [N]` | YOLOX type 6 (TDA4VM) | | classification | none | one score vector | float-input CNN with folded mean/scale (TDA4VM) | | segmentation | none | class-index mask (uint8 if the net ends in ArgMax) | no (param.yaml template from the TI zoo only) | Read `ti-edgeai-dev` first if new to the platform; it holds the rules this flow depends on (tools release compatible with the board SDK, gst-apps `SOC` key, `[H, W]` order, no-padding resize). SoC-dependent values (tools SOC, `--target-device`, quantization): `ti-edgeai-dev/references/platforms.md`. **Report evidence at each gate**: a phase is done when its check passed, not when the command returned. ## One manifest instead of repeated flags Write `model_manifest.py init` once (task, input policy, channel order, classes, meta-architecture, SoC and tools tag) and pass `--manifest` to compile, evaluators, `board_host_check.py` and `package_model.py`; contradicting flags are an error and packaging seals file hashes (`references/model-manifest.md`). Flags still work without a manifest. ## Phases and gates | # | Phase | Script / reference | Gate | |---|---|---|---| | 0 | Pre-flight: board SDK -> tools tag, tools container | `references/compile-tidl.md` | tag chosen from the compatibility table; image exists | | 1 | Get a TIDL-friendly ONNX | `check_onnx_for_tidl.py`, `add_input_preprocessing.py`, `references/export-onnx-for-tidl.md` | no ERRORs; input uint8 (normalisation folded); ONNX matches the framework model | | 2 | Compile + calibrate | `compile_tidl.py` (in container) | `Subgraph Compiled Successfully`; subgraph count and offload as expected | | 3 | Verify artifacts + accuracy | `check_artifacts_version.py`, `eval_onnx_det.py` / `eval_classification.py` | stamp OK; int8 within tolerance of float | | 4 | Package | `package_model.py` | param.yaml matches the task contract | | 5 | Deploy + smoke test + numerics | `deploy_to_board.sh`, `board_host_check.py` | app exit 0, `Offloaded Nodes N/N`, board == host outputs | | 6 | Accuracy / speed loop | `references/quantization-accuracy.md`, `ti-edgeai-profile-pipeline` | only if a gate fails | ## Phase 0 - pre-flight 1. **Board SDK -> tools tag**: `ssh root@ 'env | grep -i -E "EDGEAI|SDK"'`, then your SoC's column in `ti-edgeai-dev/references/sdk-versions.md` / TI's compatibility table (example: TDA4VM on SDK 11.0 -> `11_00_06_00`). 2. **Container** (once, ~10 min): `docker build -f scripts/Dockerfile.tidl-tools --build-arg TIDL_TAG= --build-arg SOC= -t ti-tidl-tools:- scripts/`. `run_in_container.sh` finds it through `TIDL_TAG` / `TOOLS_SOC` (defaults: `11_00_06_00`, `am68pa`). 3. Ask before changing the board's system state (services, firmware, folders you did not create). Copying a new model folder is fine. ## Phase 1 - TIDL-friendly ONNX ```bash python scripts/check_onnx_for_tidl.py model.onnx # static shapes? unsupported ops? mid-graph Cast? SiLU? ``` - Fix ERRORs (dynamic shapes -> export fixed or run `onnxsim`). Read WARNINGs: ops listed there run on ARM. - **Input convention**: the board feeds uint8 frames. If your model normalises float input (`(x-mean)*scale`), fold it in: `python3 scripts/add_input_preprocessing.py in.onnx out.onnx --mean ... --scale ...` (container; TI's own helper). Models that already start with a uint8 -> Cast, or need no normalisation, skip this. - Detectors: the ONNX should expose the raw head tensors that a TIDL meta-architecture file names. `scripts/export_yolox_tidl_det.py` does this for YOLOX (worked adapter: `references/export-yolox-det.md`); other families: `references/other-model-families.md`. - Prove the ONNX itself is right (framework output == ONNX output) before quantizing anything. ## Phase 2 - compile (container, minutes to ~40 min depending on model and frame count) ```bash cd work # holds the ONNX (+ prototxt); mounted at the same path in the container export WORKDIR=$PWD /scripts/run_in_container.sh "python3 -m onnxsim model.onnx model_sim.onnx && python3 \$SCRIPTS/compile_tidl.py \ --model model_sim.onnx --artifacts artifacts --calib --task \ [--meta model.prototxt --meta-arch-type 6] --frames 50" ``` Preprocess flags must mirror deployment: `--hw H W` (plain resize; detection/segmentation default) or `--resize S --crop C` (classification), `--channels bgr|rgb` (what the model expects). Calibration images: training distribution only, >= 50 for `accuracy_level 1`. The compile writes into a staging folder and only replaces `--artifacts` on success; an existing folder is kept as `artifacts.previous-