Creator · google-ai-edge
Last updated · Sep 4, 2026
Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, o
Creator · google-ai-edge
Last updated · Sep 4, 2026
Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, o
Creator · google-ai-edge
Last updated · Sep 4, 2026
Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, o
Creator · google-ai-edge
Last updated · Sep 4, 2026
Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, o
Sandbox only
Install targets
Codex install prompt
Install the "gpu-clean-conversion" agent skill from https://github.com/google-ai-edge/litert-samples/tree/main/skills/gpu-clean-conversion. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"google-ai-edge-gpu-clean-conversion","task":"Install gpu-clean-conversion","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Maintenance
fresh
3d since push
Risk
Risky
Permission surface may require sandboxing
GitHub quality
416
73/100 Quality · 76/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Potential broker, wallet, exchange, or real-money execution surface; sandbox and explicit approval are required
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
416 GitHub stars
Repo activity
416 stars, 116 forks
Maintenance
3d since push
License
Apache-2.0
Install
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
Agent should check
Copy prompt
Task: Use gpu-clean-conversion in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install
Install command: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
LLM text format
/api/skills/google-ai-edge-gpu-clean-conversion/install?format=text
Find alternatives
/api/skills/search?q=gpu-clean-conversion&limit=3
Agent prompt
Use gpu-clean-conversion for this task. Review https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install, then install with: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/google-ai-edge-gpu-clean-conversion
LLM text
/api/registry/manifest/google-ai-edge-gpu-clean-conversion?format=text
Install alias
/api/registry/install/google-ai-edge-gpu-clean-conversion
Recommend
/api/registry/recommend?task=Use%20gpu-clean-conversion%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
Browser automation
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO416 GitHub stars
Stars/forks activity
INFO416 stars, 116 forks; issue activity unavailable in current metadata
Recent maintenance
PASS3d since push
License clarity
PASSApache-2.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: gpu-clean-conversion description: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. ---
# GPU-clean conversion
A conversion is done when three things hold, in this order:
1. it converts, 2. every node runs on the GPU, 3. **the on-device output matches the source model.**
Step 3 is not implied by step 2. The delegate can report full residency and still return wrong numbers — that is the failure mode most of the rewrites below exist for. Never call a model done without a numerical check against CPU or PyTorch on the actual device.
## Loop
**1. Convert plain first.** No patches. This tells you what the model actually needs rather than what you assumed.
**2. Verify through the CompiledModel API before touching a device.**
```python from litert_gpu_toolkit import check_gpu_compatibility, print_report print_report(check_gpu_compatibility("model.tflite")) ```
It compiles the model for the GPU accelerator, runs every signature on random inputs, and compares the outputs against a CPU-compiled reference. A failed GPU compile names the offending op — route it through the table below. A pass with a CPU-fallback warning means some ops fell back to the CPU; treat them the same way if you need full residency. This exercises the host GPU, so step 5 on the actual device is still required.
**3. Map each symptom to a rewrite.** The rewrites live in `utilities/litert_gpu_toolkit` — plain Python, no build step. Outside this repo, clone litert-samples and put `utilities/` on `PYTHONPATH`, or vendor the directory:
| What you see | Rewrite | |---|---| | `GATHER_ND` named in the compile error | Find its source. Stride-2 slicing (`x[:, ::2]`, Focus stems, patch merging), `grid_sample`, `MaxPool` padding, bicubic `interpolate`, and reflect-mode `F.pad` all lower to it. `patch_grid_sample`, `patch_maxpool_zeropad`, `patch_interpolate`, `patch_patch_merging` | | A rank-5+ tensor named in the compile error | `PixelShuffle` (rank-6 reshape), windowed attention, `einops.rearrange`, packed-QKV attention head splits. `pixelshuffle_to_conv_transpose`, `patch_window_attention`, `patch_einops` | | `TRANSPOSE_CONV` rejected | Version skew, not a missing op. `ZeroStuffConvT1d` / `ZeroStuffConvT2d` — zero-stuff plus a plain conv, exact to ~1e-7 | | `SELECT` / `SELECT_V2` | `PReLU`, `ELU`, in-place index assignment, and `torch.where` masking. Replace with arithmetic: `x*(1-m) + v*m` | | `BROADCAST_TO` | Two cases. On a compile-time constant: an outer product or `.expand` that did not fold — bake the result as a constant at its target shape. On a **runtime** tensor the GPU delegate rejects it outright, even at rank 4 — the canonical case is GQA's `repeat_kv` (`x[:,:,None].expand(...)`, which is also rank-5, so two walls in one line). Exact rewrite: `torch.cat([x[:, i:i+1].expand(b, n_rep, s, d) for i in range(n_kv)], dim=1)` — same head order, bit-exact. Tracked upstream: google-ai-edge/LiteRT#9191 | | Masked attention wrong only on device: token 0 bit-exact, every later token wrong | Broadcast `ADD` whose LHS is a `BATCH_MATMUL` result (the `scores + mask[1,1,S,S]` idiom) silently miscomputed on older runtimes (fixed in newer; head-axis size-1 broadcast only). The signature mimics broken RoPE — tap the rope output before blaming it. Rewrites: pre-expand the mask to `[1,H,S,S]`, or materialize the BMM as a second output | | An `ADD` result that is both a graph output **and** consumed downstream comes back wrong | Output aliasing: the returned tensor holds an *operand*, not the sum — `ADD` with two runtime operands (`SUB`/`MUL` exact, `x + 1.0` exact). This is the shape of every explicit state update in a streaming/recurrent graph. Workaround: emit `acc * one` where `one` is a **runtime** input holding 1.0 — a constant folds straight back into the pattern. Tracked upstream: google-ai-edge/LiteRT#8599 | | `RELU_0_TO_1` rejected by the GPU delegate | Emitted by hard-sigmoid / `nn.Hardtanh(0,1)`. Accepted in litert 2.1.3, rejected from 2.1.5 on — a model at full residency on an older runtime hard-fails `CompiledModel` creation after an upgrade. Rewrite: `relu(x) - relu(x-1)`, exact. Tracked upstream: google-ai-edge/LiteRT#8598 | | `DIV: No support of few identical inputs` / `Expected 1 const input tensor(s)`, device only | The delegate declines an op whose two inputs are the same tensor, and ops whose inputs are all constants — together these split a perceiver-style block (softmax over a length-1 axis of a constant latent bank) into several partitions. Fixes: special-case the degenerate axis (a softmax over a length-1 axis is identically 1), or make one input non-constant. Note the sibling LayerNorm-over-a-constant pattern no longer reaches the delegate at all — the converter folds it to a single `MUL`. Tracked upstream: google-ai-edge/LiteRT#9192 | | `NHWC node rewriter not found: amax` | `x.amax(...)`/`x.max(dim)` in stable-softmax, adaptive norms, qk-norm. Rewrite channel reduce-max as `max_pool2d(x.reshape(N,1,C,H*W), kernel=(C,1))` — numerically identical — or drop the norm to 3D | | `Lowering not found: aten._fft_r2c` / `aten.complex` | `torch.stft`/`istft` and complex views have no lowering (fails before any GPU check). A DFT is a fixed linear map: windowed-DFT as `Conv1d` with the cos/sin basis baked into kernels (stride = hop), iSTFT as inverse-DFT matmul + overlap-add via zero-stuffed conv-transpose — exact. Model-selection corollary: prefer time-domain vocoder branches over iSTFT-based ones. ⚠ Library STFT-as-conv stacks (torchlibrosa-style) have **numerically mis-converted** (corr 0.83) while the op check looks clean — verify the spectrogram numerically or compute log-mel host-side | | Compiles, runs, output is wrong or NaN | The fp16 reduction family — and the trigger is the **fp16 accumulator passing 65504**, so a plain single-axis mean/sum over enough elements overflows just like variance does. `patch_safe_layernorm`, `patch_rmsnorm`, `patch_instance_norm`, `hierarchical_mean`. Caveats: at extreme magnitudes (\|x\| in the thousands) even the adaptive safe-LN form overflows when it reconstructs the large variance — the robust form stays entirely in the down-scaled domain (`xs = x/S`, normalize `xs`, never multiply the variance back by `S²`); `hierarchical_mean` is exact only for power-of-two spatial dims (for arbitrary dims, cascade `/2` avg-pools with `ceil_mode` so each stage averages ≤~49 elements). Diagnostic split: **all-zero/all-blank output = an overflow in one block; a result that starts near-correct and degrades with depth = precision compounding**, which no overflow patch (and no fp32-precision flag) fixes | | Head outputs exactly zero | RMSNorm `Σx²` overflowed fp16 to `inf`. `patch_rmsnorm` |
**4. Re-convert and re-verify.** Repeat 2–3 until the check passes.
**5. Verify on device, numerically.** Run the same input through the source model and the on-device graph and compare. Correlation on the output tensor, plus a task-level check where one exists (argmax match, IoU, mask foreground count). Record the device and the residency line (`N/N` nodes) alongside the number — a result without both is not reproducible.
If residency is full and the output is wrong, bisect by intermediate: dump tensors at block boundaries and find the first one that diverges. The cause is usually a reduction, not the op you suspect.
## Watch for
- **fp16 is used even for an fp32 graph.** The delegate reduces in fp16 regardless of tensor dtype. Anything that sums many large values — variance, `Σx²`, multi-axis mean — is a candidate. - **Approximation choices are not free.** Substituting a GELU or Swish approximation changes numerics. A head with a wide output range can lose real accuracy to the coarser form. - **`inplace=True` activations mutate a residual.** `out = self.act(x)` then `x + out` adds `relu(x)`, not `x`. Swapping the activation silently changes the math. - **A silent CPU fallback looks like success.** If the GPU output is bit-identical to CPU fp32, it probably did not run on the GPU. Genuine fp16 execution drifts in the last digits. - **Reflect padding routes through `F.pad` even at `padding=0`** — every conv with `padding_mode='reflect'` is affected, not just padded ones. Slice+concat reflect pads are cheap for small pad widths. - **Channel-attention through a `Linear` confuses layout handling at convert time** (`tfl.mul operands don't have broadcast-compatible shapes`) — express channel attention as a 1×1 conv. - **`.chunk()` lowers to `SPLIT`** (GPU-rejected) — slice directly instead; bit-exact. - **Constant folding can explode the file**: a frozen-param × constant product materializes at full size per block (fp16 casting skips non-weight constants, so it cannot rescue it). Feed one factor as a runtime input — param × input never folds.
## Output layout
Ship the result as a model recipe, not a one-off script:
``` models/<family>/<model>/ README.md what the model is, what was produced converted/ README.md environments, pipeline, verification results export_*.py one script per graph dump_*_ref.py reference dumps from the source model verify_*.py parity check against those references ```
Keep the export, the reference dump, and the verification separate so each can be re-run alone. State the verification numbers in the README with the device they came from. Weights are not committed.
The tree above is this repo's convention. In your own project the directory names matter less than the split — keep the same three separately-runnable pieces wherever your models live.
Do the model first. A demo can follow, and reusable inference code matters more in it than a full UI.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for gpu-clean-conversion, ready for a manual X post.
gpu-clean-conversion: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via th... 416 stars https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x
Listing + install path for gpu-clean-conversion: https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x Install: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to google-ai-edge but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion/audit)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)google-ai-edge
@google-ai-edge
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "gpu-clean-conversion" agent skill from https://github.com/google-ai-edge/litert-samples/tree/main/skills/gpu-clean-conversion. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"google-ai-edge-gpu-clean-conversion","task":"Install gpu-clean-conversion","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Maintenance
fresh
3d since push
Risk
Risky
Permission surface may require sandboxing
GitHub quality
416
73/100 Quality · 76/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Potential broker, wallet, exchange, or real-money execution surface; sandbox and explicit approval are required
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
416 GitHub stars
Repo activity
416 stars, 116 forks
Maintenance
3d since push
License
Apache-2.0
Install
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
Agent should check
Copy prompt
Task: Use gpu-clean-conversion in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install
Install command: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
LLM text format
/api/skills/google-ai-edge-gpu-clean-conversion/install?format=text
Find alternatives
/api/skills/search?q=gpu-clean-conversion&limit=3
Agent prompt
Use gpu-clean-conversion for this task. Review https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install, then install with: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/google-ai-edge-gpu-clean-conversion
LLM text
/api/registry/manifest/google-ai-edge-gpu-clean-conversion?format=text
Install alias
/api/registry/install/google-ai-edge-gpu-clean-conversion
Recommend
/api/registry/recommend?task=Use%20gpu-clean-conversion%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
Browser automation
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO416 GitHub stars
Stars/forks activity
INFO416 stars, 116 forks; issue activity unavailable in current metadata
Recent maintenance
PASS3d since push
License clarity
PASSApache-2.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: gpu-clean-conversion description: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. ---
# GPU-clean conversion
A conversion is done when three things hold, in this order:
1. it converts, 2. every node runs on the GPU, 3. **the on-device output matches the source model.**
Step 3 is not implied by step 2. The delegate can report full residency and still return wrong numbers — that is the failure mode most of the rewrites below exist for. Never call a model done without a numerical check against CPU or PyTorch on the actual device.
## Loop
**1. Convert plain first.** No patches. This tells you what the model actually needs rather than what you assumed.
**2. Verify through the CompiledModel API before touching a device.**
```python from litert_gpu_toolkit import check_gpu_compatibility, print_report print_report(check_gpu_compatibility("model.tflite")) ```
It compiles the model for the GPU accelerator, runs every signature on random inputs, and compares the outputs against a CPU-compiled reference. A failed GPU compile names the offending op — route it through the table below. A pass with a CPU-fallback warning means some ops fell back to the CPU; treat them the same way if you need full residency. This exercises the host GPU, so step 5 on the actual device is still required.
**3. Map each symptom to a rewrite.** The rewrites live in `utilities/litert_gpu_toolkit` — plain Python, no build step. Outside this repo, clone litert-samples and put `utilities/` on `PYTHONPATH`, or vendor the directory:
| What you see | Rewrite | |---|---| | `GATHER_ND` named in the compile error | Find its source. Stride-2 slicing (`x[:, ::2]`, Focus stems, patch merging), `grid_sample`, `MaxPool` padding, bicubic `interpolate`, and reflect-mode `F.pad` all lower to it. `patch_grid_sample`, `patch_maxpool_zeropad`, `patch_interpolate`, `patch_patch_merging` | | A rank-5+ tensor named in the compile error | `PixelShuffle` (rank-6 reshape), windowed attention, `einops.rearrange`, packed-QKV attention head splits. `pixelshuffle_to_conv_transpose`, `patch_window_attention`, `patch_einops` | | `TRANSPOSE_CONV` rejected | Version skew, not a missing op. `ZeroStuffConvT1d` / `ZeroStuffConvT2d` — zero-stuff plus a plain conv, exact to ~1e-7 | | `SELECT` / `SELECT_V2` | `PReLU`, `ELU`, in-place index assignment, and `torch.where` masking. Replace with arithmetic: `x*(1-m) + v*m` | | `BROADCAST_TO` | Two cases. On a compile-time constant: an outer product or `.expand` that did not fold — bake the result as a constant at its target shape. On a **runtime** tensor the GPU delegate rejects it outright, even at rank 4 — the canonical case is GQA's `repeat_kv` (`x[:,:,None].expand(...)`, which is also rank-5, so two walls in one line). Exact rewrite: `torch.cat([x[:, i:i+1].expand(b, n_rep, s, d) for i in range(n_kv)], dim=1)` — same head order, bit-exact. Tracked upstream: google-ai-edge/LiteRT#9191 | | Masked attention wrong only on device: token 0 bit-exact, every later token wrong | Broadcast `ADD` whose LHS is a `BATCH_MATMUL` result (the `scores + mask[1,1,S,S]` idiom) silently miscomputed on older runtimes (fixed in newer; head-axis size-1 broadcast only). The signature mimics broken RoPE — tap the rope output before blaming it. Rewrites: pre-expand the mask to `[1,H,S,S]`, or materialize the BMM as a second output | | An `ADD` result that is both a graph output **and** consumed downstream comes back wrong | Output aliasing: the returned tensor holds an *operand*, not the sum — `ADD` with two runtime operands (`SUB`/`MUL` exact, `x + 1.0` exact). This is the shape of every explicit state update in a streaming/recurrent graph. Workaround: emit `acc * one` where `one` is a **runtime** input holding 1.0 — a constant folds straight back into the pattern. Tracked upstream: google-ai-edge/LiteRT#8599 | | `RELU_0_TO_1` rejected by the GPU delegate | Emitted by hard-sigmoid / `nn.Hardtanh(0,1)`. Accepted in litert 2.1.3, rejected from 2.1.5 on — a model at full residency on an older runtime hard-fails `CompiledModel` creation after an upgrade. Rewrite: `relu(x) - relu(x-1)`, exact. Tracked upstream: google-ai-edge/LiteRT#8598 | | `DIV: No support of few identical inputs` / `Expected 1 const input tensor(s)`, device only | The delegate declines an op whose two inputs are the same tensor, and ops whose inputs are all constants — together these split a perceiver-style block (softmax over a length-1 axis of a constant latent bank) into several partitions. Fixes: special-case the degenerate axis (a softmax over a length-1 axis is identically 1), or make one input non-constant. Note the sibling LayerNorm-over-a-constant pattern no longer reaches the delegate at all — the converter folds it to a single `MUL`. Tracked upstream: google-ai-edge/LiteRT#9192 | | `NHWC node rewriter not found: amax` | `x.amax(...)`/`x.max(dim)` in stable-softmax, adaptive norms, qk-norm. Rewrite channel reduce-max as `max_pool2d(x.reshape(N,1,C,H*W), kernel=(C,1))` — numerically identical — or drop the norm to 3D | | `Lowering not found: aten._fft_r2c` / `aten.complex` | `torch.stft`/`istft` and complex views have no lowering (fails before any GPU check). A DFT is a fixed linear map: windowed-DFT as `Conv1d` with the cos/sin basis baked into kernels (stride = hop), iSTFT as inverse-DFT matmul + overlap-add via zero-stuffed conv-transpose — exact. Model-selection corollary: prefer time-domain vocoder branches over iSTFT-based ones. ⚠ Library STFT-as-conv stacks (torchlibrosa-style) have **numerically mis-converted** (corr 0.83) while the op check looks clean — verify the spectrogram numerically or compute log-mel host-side | | Compiles, runs, output is wrong or NaN | The fp16 reduction family — and the trigger is the **fp16 accumulator passing 65504**, so a plain single-axis mean/sum over enough elements overflows just like variance does. `patch_safe_layernorm`, `patch_rmsnorm`, `patch_instance_norm`, `hierarchical_mean`. Caveats: at extreme magnitudes (\|x\| in the thousands) even the adaptive safe-LN form overflows when it reconstructs the large variance — the robust form stays entirely in the down-scaled domain (`xs = x/S`, normalize `xs`, never multiply the variance back by `S²`); `hierarchical_mean` is exact only for power-of-two spatial dims (for arbitrary dims, cascade `/2` avg-pools with `ceil_mode` so each stage averages ≤~49 elements). Diagnostic split: **all-zero/all-blank output = an overflow in one block; a result that starts near-correct and degrades with depth = precision compounding**, which no overflow patch (and no fp32-precision flag) fixes | | Head outputs exactly zero | RMSNorm `Σx²` overflowed fp16 to `inf`. `patch_rmsnorm` |
**4. Re-convert and re-verify.** Repeat 2–3 until the check passes.
**5. Verify on device, numerically.** Run the same input through the source model and the on-device graph and compare. Correlation on the output tensor, plus a task-level check where one exists (argmax match, IoU, mask foreground count). Record the device and the residency line (`N/N` nodes) alongside the number — a result without both is not reproducible.
If residency is full and the output is wrong, bisect by intermediate: dump tensors at block boundaries and find the first one that diverges. The cause is usually a reduction, not the op you suspect.
## Watch for
- **fp16 is used even for an fp32 graph.** The delegate reduces in fp16 regardless of tensor dtype. Anything that sums many large values — variance, `Σx²`, multi-axis mean — is a candidate. - **Approximation choices are not free.** Substituting a GELU or Swish approximation changes numerics. A head with a wide output range can lose real accuracy to the coarser form. - **`inplace=True` activations mutate a residual.** `out = self.act(x)` then `x + out` adds `relu(x)`, not `x`. Swapping the activation silently changes the math. - **A silent CPU fallback looks like success.** If the GPU output is bit-identical to CPU fp32, it probably did not run on the GPU. Genuine fp16 execution drifts in the last digits. - **Reflect padding routes through `F.pad` even at `padding=0`** — every conv with `padding_mode='reflect'` is affected, not just padded ones. Slice+concat reflect pads are cheap for small pad widths. - **Channel-attention through a `Linear` confuses layout handling at convert time** (`tfl.mul operands don't have broadcast-compatible shapes`) — express channel attention as a 1×1 conv. - **`.chunk()` lowers to `SPLIT`** (GPU-rejected) — slice directly instead; bit-exact. - **Constant folding can explode the file**: a frozen-param × constant product materializes at full size per block (fp16 casting skips non-weight constants, so it cannot rescue it). Feed one factor as a runtime input — param × input never folds.
## Output layout
Ship the result as a model recipe, not a one-off script:
``` models/<family>/<model>/ README.md what the model is, what was produced converted/ README.md environments, pipeline, verification results export_*.py one script per graph dump_*_ref.py reference dumps from the source model verify_*.py parity check against those references ```
Keep the export, the reference dump, and the verification separate so each can be re-run alone. State the verification numbers in the README with the device they came from. Weights are not committed.
The tree above is this repo's convention. In your own project the directory names matter less than the split — keep the same three separately-runnable pieces wherever your models live.
Do the model first. A demo can follow, and reusable inference code matters more in it than a full UI.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for gpu-clean-conversion, ready for a manual X post.
gpu-clean-conversion: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via th... 416 stars https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x
Listing + install path for gpu-clean-conversion: https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x Install: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to google-ai-edge but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion/audit)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)google-ai-edge
@google-ai-edge
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "gpu-clean-conversion" agent skill from https://github.com/google-ai-edge/litert-samples/tree/main/skills/gpu-clean-conversion. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"google-ai-edge-gpu-clean-conversion","task":"Install gpu-clean-conversion","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Maintenance
fresh
3d since push
Risk
Risky
Permission surface may require sandboxing
GitHub quality
416
73/100 Quality · 76/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Potential broker, wallet, exchange, or real-money execution surface; sandbox and explicit approval are required
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
416 GitHub stars
Repo activity
416 stars, 116 forks
Maintenance
3d since push
License
Apache-2.0
Install
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
Agent should check
Copy prompt
Task: Use gpu-clean-conversion in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install
Install command: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
LLM text format
/api/skills/google-ai-edge-gpu-clean-conversion/install?format=text
Find alternatives
/api/skills/search?q=gpu-clean-conversion&limit=3
Agent prompt
Use gpu-clean-conversion for this task. Review https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install, then install with: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/google-ai-edge-gpu-clean-conversion
LLM text
/api/registry/manifest/google-ai-edge-gpu-clean-conversion?format=text
Install alias
/api/registry/install/google-ai-edge-gpu-clean-conversion
Recommend
/api/registry/recommend?task=Use%20gpu-clean-conversion%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
Browser automation
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO416 GitHub stars
Stars/forks activity
INFO416 stars, 116 forks; issue activity unavailable in current metadata
Recent maintenance
PASS3d since push
License clarity
PASSApache-2.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: gpu-clean-conversion description: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. ---
# GPU-clean conversion
A conversion is done when three things hold, in this order:
1. it converts, 2. every node runs on the GPU, 3. **the on-device output matches the source model.**
Step 3 is not implied by step 2. The delegate can report full residency and still return wrong numbers — that is the failure mode most of the rewrites below exist for. Never call a model done without a numerical check against CPU or PyTorch on the actual device.
## Loop
**1. Convert plain first.** No patches. This tells you what the model actually needs rather than what you assumed.
**2. Verify through the CompiledModel API before touching a device.**
```python from litert_gpu_toolkit import check_gpu_compatibility, print_report print_report(check_gpu_compatibility("model.tflite")) ```
It compiles the model for the GPU accelerator, runs every signature on random inputs, and compares the outputs against a CPU-compiled reference. A failed GPU compile names the offending op — route it through the table below. A pass with a CPU-fallback warning means some ops fell back to the CPU; treat them the same way if you need full residency. This exercises the host GPU, so step 5 on the actual device is still required.
**3. Map each symptom to a rewrite.** The rewrites live in `utilities/litert_gpu_toolkit` — plain Python, no build step. Outside this repo, clone litert-samples and put `utilities/` on `PYTHONPATH`, or vendor the directory:
| What you see | Rewrite | |---|---| | `GATHER_ND` named in the compile error | Find its source. Stride-2 slicing (`x[:, ::2]`, Focus stems, patch merging), `grid_sample`, `MaxPool` padding, bicubic `interpolate`, and reflect-mode `F.pad` all lower to it. `patch_grid_sample`, `patch_maxpool_zeropad`, `patch_interpolate`, `patch_patch_merging` | | A rank-5+ tensor named in the compile error | `PixelShuffle` (rank-6 reshape), windowed attention, `einops.rearrange`, packed-QKV attention head splits. `pixelshuffle_to_conv_transpose`, `patch_window_attention`, `patch_einops` | | `TRANSPOSE_CONV` rejected | Version skew, not a missing op. `ZeroStuffConvT1d` / `ZeroStuffConvT2d` — zero-stuff plus a plain conv, exact to ~1e-7 | | `SELECT` / `SELECT_V2` | `PReLU`, `ELU`, in-place index assignment, and `torch.where` masking. Replace with arithmetic: `x*(1-m) + v*m` | | `BROADCAST_TO` | Two cases. On a compile-time constant: an outer product or `.expand` that did not fold — bake the result as a constant at its target shape. On a **runtime** tensor the GPU delegate rejects it outright, even at rank 4 — the canonical case is GQA's `repeat_kv` (`x[:,:,None].expand(...)`, which is also rank-5, so two walls in one line). Exact rewrite: `torch.cat([x[:, i:i+1].expand(b, n_rep, s, d) for i in range(n_kv)], dim=1)` — same head order, bit-exact. Tracked upstream: google-ai-edge/LiteRT#9191 | | Masked attention wrong only on device: token 0 bit-exact, every later token wrong | Broadcast `ADD` whose LHS is a `BATCH_MATMUL` result (the `scores + mask[1,1,S,S]` idiom) silently miscomputed on older runtimes (fixed in newer; head-axis size-1 broadcast only). The signature mimics broken RoPE — tap the rope output before blaming it. Rewrites: pre-expand the mask to `[1,H,S,S]`, or materialize the BMM as a second output | | An `ADD` result that is both a graph output **and** consumed downstream comes back wrong | Output aliasing: the returned tensor holds an *operand*, not the sum — `ADD` with two runtime operands (`SUB`/`MUL` exact, `x + 1.0` exact). This is the shape of every explicit state update in a streaming/recurrent graph. Workaround: emit `acc * one` where `one` is a **runtime** input holding 1.0 — a constant folds straight back into the pattern. Tracked upstream: google-ai-edge/LiteRT#8599 | | `RELU_0_TO_1` rejected by the GPU delegate | Emitted by hard-sigmoid / `nn.Hardtanh(0,1)`. Accepted in litert 2.1.3, rejected from 2.1.5 on — a model at full residency on an older runtime hard-fails `CompiledModel` creation after an upgrade. Rewrite: `relu(x) - relu(x-1)`, exact. Tracked upstream: google-ai-edge/LiteRT#8598 | | `DIV: No support of few identical inputs` / `Expected 1 const input tensor(s)`, device only | The delegate declines an op whose two inputs are the same tensor, and ops whose inputs are all constants — together these split a perceiver-style block (softmax over a length-1 axis of a constant latent bank) into several partitions. Fixes: special-case the degenerate axis (a softmax over a length-1 axis is identically 1), or make one input non-constant. Note the sibling LayerNorm-over-a-constant pattern no longer reaches the delegate at all — the converter folds it to a single `MUL`. Tracked upstream: google-ai-edge/LiteRT#9192 | | `NHWC node rewriter not found: amax` | `x.amax(...)`/`x.max(dim)` in stable-softmax, adaptive norms, qk-norm. Rewrite channel reduce-max as `max_pool2d(x.reshape(N,1,C,H*W), kernel=(C,1))` — numerically identical — or drop the norm to 3D | | `Lowering not found: aten._fft_r2c` / `aten.complex` | `torch.stft`/`istft` and complex views have no lowering (fails before any GPU check). A DFT is a fixed linear map: windowed-DFT as `Conv1d` with the cos/sin basis baked into kernels (stride = hop), iSTFT as inverse-DFT matmul + overlap-add via zero-stuffed conv-transpose — exact. Model-selection corollary: prefer time-domain vocoder branches over iSTFT-based ones. ⚠ Library STFT-as-conv stacks (torchlibrosa-style) have **numerically mis-converted** (corr 0.83) while the op check looks clean — verify the spectrogram numerically or compute log-mel host-side | | Compiles, runs, output is wrong or NaN | The fp16 reduction family — and the trigger is the **fp16 accumulator passing 65504**, so a plain single-axis mean/sum over enough elements overflows just like variance does. `patch_safe_layernorm`, `patch_rmsnorm`, `patch_instance_norm`, `hierarchical_mean`. Caveats: at extreme magnitudes (\|x\| in the thousands) even the adaptive safe-LN form overflows when it reconstructs the large variance — the robust form stays entirely in the down-scaled domain (`xs = x/S`, normalize `xs`, never multiply the variance back by `S²`); `hierarchical_mean` is exact only for power-of-two spatial dims (for arbitrary dims, cascade `/2` avg-pools with `ceil_mode` so each stage averages ≤~49 elements). Diagnostic split: **all-zero/all-blank output = an overflow in one block; a result that starts near-correct and degrades with depth = precision compounding**, which no overflow patch (and no fp32-precision flag) fixes | | Head outputs exactly zero | RMSNorm `Σx²` overflowed fp16 to `inf`. `patch_rmsnorm` |
**4. Re-convert and re-verify.** Repeat 2–3 until the check passes.
**5. Verify on device, numerically.** Run the same input through the source model and the on-device graph and compare. Correlation on the output tensor, plus a task-level check where one exists (argmax match, IoU, mask foreground count). Record the device and the residency line (`N/N` nodes) alongside the number — a result without both is not reproducible.
If residency is full and the output is wrong, bisect by intermediate: dump tensors at block boundaries and find the first one that diverges. The cause is usually a reduction, not the op you suspect.
## Watch for
- **fp16 is used even for an fp32 graph.** The delegate reduces in fp16 regardless of tensor dtype. Anything that sums many large values — variance, `Σx²`, multi-axis mean — is a candidate. - **Approximation choices are not free.** Substituting a GELU or Swish approximation changes numerics. A head with a wide output range can lose real accuracy to the coarser form. - **`inplace=True` activations mutate a residual.** `out = self.act(x)` then `x + out` adds `relu(x)`, not `x`. Swapping the activation silently changes the math. - **A silent CPU fallback looks like success.** If the GPU output is bit-identical to CPU fp32, it probably did not run on the GPU. Genuine fp16 execution drifts in the last digits. - **Reflect padding routes through `F.pad` even at `padding=0`** — every conv with `padding_mode='reflect'` is affected, not just padded ones. Slice+concat reflect pads are cheap for small pad widths. - **Channel-attention through a `Linear` confuses layout handling at convert time** (`tfl.mul operands don't have broadcast-compatible shapes`) — express channel attention as a 1×1 conv. - **`.chunk()` lowers to `SPLIT`** (GPU-rejected) — slice directly instead; bit-exact. - **Constant folding can explode the file**: a frozen-param × constant product materializes at full size per block (fp16 casting skips non-weight constants, so it cannot rescue it). Feed one factor as a runtime input — param × input never folds.
## Output layout
Ship the result as a model recipe, not a one-off script:
``` models/<family>/<model>/ README.md what the model is, what was produced converted/ README.md environments, pipeline, verification results export_*.py one script per graph dump_*_ref.py reference dumps from the source model verify_*.py parity check against those references ```
Keep the export, the reference dump, and the verification separate so each can be re-run alone. State the verification numbers in the README with the device they came from. Weights are not committed.
The tree above is this repo's convention. In your own project the directory names matter less than the split — keep the same three separately-runnable pieces wherever your models live.
Do the model first. A demo can follow, and reusable inference code matters more in it than a full UI.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for gpu-clean-conversion, ready for a manual X post.
gpu-clean-conversion: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via th... 416 stars https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x
Listing + install path for gpu-clean-conversion: https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x Install: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to google-ai-edge but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion/audit)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)google-ai-edge
@google-ai-edge
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "gpu-clean-conversion" agent skill from https://github.com/google-ai-edge/litert-samples/tree/main/skills/gpu-clean-conversion. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"google-ai-edge-gpu-clean-conversion","task":"Install gpu-clean-conversion","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Maintenance
fresh
3d since push
Risk
Risky
Permission surface may require sandboxing
GitHub quality
416
73/100 Quality · 76/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Potential broker, wallet, exchange, or real-money execution surface; sandbox and explicit approval are required
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
416 GitHub stars
Repo activity
416 stars, 116 forks
Maintenance
3d since push
License
Apache-2.0
Install
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
Agent should check
Copy prompt
Task: Use gpu-clean-conversion in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20gpu-clean-conversion%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install
Install command: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/google-ai-edge-gpu-clean-conversion/install
LLM text format
/api/skills/google-ai-edge-gpu-clean-conversion/install?format=text
Find alternatives
/api/skills/search?q=gpu-clean-conversion&limit=3
Agent prompt
Use gpu-clean-conversion for this task. Review https://www.openagentskill.com/api/skills/google-ai-edge-gpu-clean-conversion/install, then install with: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversionRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/google-ai-edge-gpu-clean-conversion
LLM text
/api/registry/manifest/google-ai-edge-gpu-clean-conversion?format=text
Install alias
/api/registry/install/google-ai-edge-gpu-clean-conversion
Recommend
/api/registry/recommend?task=Use%20gpu-clean-conversion%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
Browser automation
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO416 GitHub stars
Stars/forks activity
INFO416 stars, 116 forks; issue activity unavailable in current metadata
Recent maintenance
PASS3d since push
License clarity
PASSApache-2.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: gpu-clean-conversion description: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via the CompiledModel API with verified-correct output, and lay it out as a model recipe. Use when converting a new model, or when a converted model is rejected by the GPU, falls back to CPU, or returns wrong numbers on device. ---
# GPU-clean conversion
A conversion is done when three things hold, in this order:
1. it converts, 2. every node runs on the GPU, 3. **the on-device output matches the source model.**
Step 3 is not implied by step 2. The delegate can report full residency and still return wrong numbers — that is the failure mode most of the rewrites below exist for. Never call a model done without a numerical check against CPU or PyTorch on the actual device.
## Loop
**1. Convert plain first.** No patches. This tells you what the model actually needs rather than what you assumed.
**2. Verify through the CompiledModel API before touching a device.**
```python from litert_gpu_toolkit import check_gpu_compatibility, print_report print_report(check_gpu_compatibility("model.tflite")) ```
It compiles the model for the GPU accelerator, runs every signature on random inputs, and compares the outputs against a CPU-compiled reference. A failed GPU compile names the offending op — route it through the table below. A pass with a CPU-fallback warning means some ops fell back to the CPU; treat them the same way if you need full residency. This exercises the host GPU, so step 5 on the actual device is still required.
**3. Map each symptom to a rewrite.** The rewrites live in `utilities/litert_gpu_toolkit` — plain Python, no build step. Outside this repo, clone litert-samples and put `utilities/` on `PYTHONPATH`, or vendor the directory:
| What you see | Rewrite | |---|---| | `GATHER_ND` named in the compile error | Find its source. Stride-2 slicing (`x[:, ::2]`, Focus stems, patch merging), `grid_sample`, `MaxPool` padding, bicubic `interpolate`, and reflect-mode `F.pad` all lower to it. `patch_grid_sample`, `patch_maxpool_zeropad`, `patch_interpolate`, `patch_patch_merging` | | A rank-5+ tensor named in the compile error | `PixelShuffle` (rank-6 reshape), windowed attention, `einops.rearrange`, packed-QKV attention head splits. `pixelshuffle_to_conv_transpose`, `patch_window_attention`, `patch_einops` | | `TRANSPOSE_CONV` rejected | Version skew, not a missing op. `ZeroStuffConvT1d` / `ZeroStuffConvT2d` — zero-stuff plus a plain conv, exact to ~1e-7 | | `SELECT` / `SELECT_V2` | `PReLU`, `ELU`, in-place index assignment, and `torch.where` masking. Replace with arithmetic: `x*(1-m) + v*m` | | `BROADCAST_TO` | Two cases. On a compile-time constant: an outer product or `.expand` that did not fold — bake the result as a constant at its target shape. On a **runtime** tensor the GPU delegate rejects it outright, even at rank 4 — the canonical case is GQA's `repeat_kv` (`x[:,:,None].expand(...)`, which is also rank-5, so two walls in one line). Exact rewrite: `torch.cat([x[:, i:i+1].expand(b, n_rep, s, d) for i in range(n_kv)], dim=1)` — same head order, bit-exact. Tracked upstream: google-ai-edge/LiteRT#9191 | | Masked attention wrong only on device: token 0 bit-exact, every later token wrong | Broadcast `ADD` whose LHS is a `BATCH_MATMUL` result (the `scores + mask[1,1,S,S]` idiom) silently miscomputed on older runtimes (fixed in newer; head-axis size-1 broadcast only). The signature mimics broken RoPE — tap the rope output before blaming it. Rewrites: pre-expand the mask to `[1,H,S,S]`, or materialize the BMM as a second output | | An `ADD` result that is both a graph output **and** consumed downstream comes back wrong | Output aliasing: the returned tensor holds an *operand*, not the sum — `ADD` with two runtime operands (`SUB`/`MUL` exact, `x + 1.0` exact). This is the shape of every explicit state update in a streaming/recurrent graph. Workaround: emit `acc * one` where `one` is a **runtime** input holding 1.0 — a constant folds straight back into the pattern. Tracked upstream: google-ai-edge/LiteRT#8599 | | `RELU_0_TO_1` rejected by the GPU delegate | Emitted by hard-sigmoid / `nn.Hardtanh(0,1)`. Accepted in litert 2.1.3, rejected from 2.1.5 on — a model at full residency on an older runtime hard-fails `CompiledModel` creation after an upgrade. Rewrite: `relu(x) - relu(x-1)`, exact. Tracked upstream: google-ai-edge/LiteRT#8598 | | `DIV: No support of few identical inputs` / `Expected 1 const input tensor(s)`, device only | The delegate declines an op whose two inputs are the same tensor, and ops whose inputs are all constants — together these split a perceiver-style block (softmax over a length-1 axis of a constant latent bank) into several partitions. Fixes: special-case the degenerate axis (a softmax over a length-1 axis is identically 1), or make one input non-constant. Note the sibling LayerNorm-over-a-constant pattern no longer reaches the delegate at all — the converter folds it to a single `MUL`. Tracked upstream: google-ai-edge/LiteRT#9192 | | `NHWC node rewriter not found: amax` | `x.amax(...)`/`x.max(dim)` in stable-softmax, adaptive norms, qk-norm. Rewrite channel reduce-max as `max_pool2d(x.reshape(N,1,C,H*W), kernel=(C,1))` — numerically identical — or drop the norm to 3D | | `Lowering not found: aten._fft_r2c` / `aten.complex` | `torch.stft`/`istft` and complex views have no lowering (fails before any GPU check). A DFT is a fixed linear map: windowed-DFT as `Conv1d` with the cos/sin basis baked into kernels (stride = hop), iSTFT as inverse-DFT matmul + overlap-add via zero-stuffed conv-transpose — exact. Model-selection corollary: prefer time-domain vocoder branches over iSTFT-based ones. ⚠ Library STFT-as-conv stacks (torchlibrosa-style) have **numerically mis-converted** (corr 0.83) while the op check looks clean — verify the spectrogram numerically or compute log-mel host-side | | Compiles, runs, output is wrong or NaN | The fp16 reduction family — and the trigger is the **fp16 accumulator passing 65504**, so a plain single-axis mean/sum over enough elements overflows just like variance does. `patch_safe_layernorm`, `patch_rmsnorm`, `patch_instance_norm`, `hierarchical_mean`. Caveats: at extreme magnitudes (\|x\| in the thousands) even the adaptive safe-LN form overflows when it reconstructs the large variance — the robust form stays entirely in the down-scaled domain (`xs = x/S`, normalize `xs`, never multiply the variance back by `S²`); `hierarchical_mean` is exact only for power-of-two spatial dims (for arbitrary dims, cascade `/2` avg-pools with `ceil_mode` so each stage averages ≤~49 elements). Diagnostic split: **all-zero/all-blank output = an overflow in one block; a result that starts near-correct and degrades with depth = precision compounding**, which no overflow patch (and no fp32-precision flag) fixes | | Head outputs exactly zero | RMSNorm `Σx²` overflowed fp16 to `inf`. `patch_rmsnorm` |
**4. Re-convert and re-verify.** Repeat 2–3 until the check passes.
**5. Verify on device, numerically.** Run the same input through the source model and the on-device graph and compare. Correlation on the output tensor, plus a task-level check where one exists (argmax match, IoU, mask foreground count). Record the device and the residency line (`N/N` nodes) alongside the number — a result without both is not reproducible.
If residency is full and the output is wrong, bisect by intermediate: dump tensors at block boundaries and find the first one that diverges. The cause is usually a reduction, not the op you suspect.
## Watch for
- **fp16 is used even for an fp32 graph.** The delegate reduces in fp16 regardless of tensor dtype. Anything that sums many large values — variance, `Σx²`, multi-axis mean — is a candidate. - **Approximation choices are not free.** Substituting a GELU or Swish approximation changes numerics. A head with a wide output range can lose real accuracy to the coarser form. - **`inplace=True` activations mutate a residual.** `out = self.act(x)` then `x + out` adds `relu(x)`, not `x`. Swapping the activation silently changes the math. - **A silent CPU fallback looks like success.** If the GPU output is bit-identical to CPU fp32, it probably did not run on the GPU. Genuine fp16 execution drifts in the last digits. - **Reflect padding routes through `F.pad` even at `padding=0`** — every conv with `padding_mode='reflect'` is affected, not just padded ones. Slice+concat reflect pads are cheap for small pad widths. - **Channel-attention through a `Linear` confuses layout handling at convert time** (`tfl.mul operands don't have broadcast-compatible shapes`) — express channel attention as a 1×1 conv. - **`.chunk()` lowers to `SPLIT`** (GPU-rejected) — slice directly instead; bit-exact. - **Constant folding can explode the file**: a frozen-param × constant product materializes at full size per block (fp16 casting skips non-weight constants, so it cannot rescue it). Feed one factor as a runtime input — param × input never folds.
## Output layout
Ship the result as a model recipe, not a one-off script:
``` models/<family>/<model>/ README.md what the model is, what was produced converted/ README.md environments, pipeline, verification results export_*.py one script per graph dump_*_ref.py reference dumps from the source model verify_*.py parity check against those references ```
Keep the export, the reference dump, and the verification separate so each can be re-run alone. State the verification numbers in the README with the device they came from. Weights are not committed.
The tree above is this repo's convention. In your own project the directory names matter less than the split — keep the same three separately-runnable pieces wherever your models live.
Do the model first. A demo can follow, and reusable inference code matters more in it than a full UI.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for gpu-clean-conversion, ready for a manual X post.
gpu-clean-conversion: Convert a PyTorch or Hugging Face model into a LiteRT model that runs fully on the GPU via th... 416 stars https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x
Listing + install path for gpu-clean-conversion: https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=x Install: npx skills add google-ai-edge/litert-samples --skill gpu-clean-conversion
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to google-ai-edge but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion/audit)
[](https://www.openagentskill.com/skills/google-ai-edge-gpu-clean-conversion?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)google-ai-edge
@google-ai-edge
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsPermission surface
secrets or environment access, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness