暂未收录效果图
查看技能说明VLM Grounder
InternRobotics
[CoRL 2024] VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
16 Skills
搜索结果: 16
暂未收录效果图
查看技能说明InternRobotics
[CoRL 2024] VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
暂未收录效果图
查看技能说明vlm-run
A hub for various industry-specific schemas to be used with VLMs.
暂未收录效果图
查看技能说明modelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
暂未收录效果图
查看技能说明zapdos-labs
暂未收录效果图
查看技能说明zai-org
暂未收录效果图
查看技能说明yzhao062
Anomaly detection related books, papers, videos, and toolboxes. Last update late 2025 for LLM and VLM works!
暂未收录效果图
查看技能说明Achno
A tool to convert a Wallpaper's color scheme / palette, OCR with VLM's Traditional & Hybrid, Image Compression ,color palette extraction, image upsacling with Adversaria…
暂未收录效果图
查看技能说明SharpAI
Open-Source AI Camera Skills Platform, AI NVR & CCTV Surveillance. Local VLM video analysis with Qwen, DeepSeek, SmolVLM, LLaVA, YOLO26. LLM-powered agentic security cam…
暂未收录效果图
查看技能说明wanshuiyin
Turn a refined research proposal or method idea into a detailed, claim-driven experiment roadmap. Use after `research-refine`, or when the user asks for a detailed exper…
暂未收录效果图
查看技能说明eggbrid2
Open Android AI agent runtime for phone control, app automation, VLM screen reading, skill routing, mini apps, and Mihomo VPN workflows.
暂未收录效果图
查看技能说明zenstory-ai
对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验。work_dir 已包含 agent_narration_brief.md 与 vlm_analysis.json 时使用。适用于故事方向、片段选择、画面/原声/旁白分工、 解说写作与复核。输入 work_dir 中的理解索引;输出 recap_story_plan.j…
暂未收录效果图
查看技能说明google-ai-edge
Convert a Hugging Face LLM or vision-language model checkpoint into a .litertlm bundle that runs on the LiteRT-LM runtime with verified quality - classify the architectu…
暂未收录效果图
查看技能说明TianLin0509
Turn a slide mockup image into an editable PPTX: native text + semantic draggable sprites + inpainted background. Agent-as-VLM skill for Claude Code / Codex / any CLI ag…
暂未收录效果图
查看技能说明graph-robots
Lightweight one-shot 3D object perception. Runs Grounding-DINO broad detection, a SINGLE VLM set-of-marks letter pick over the labeled boxes, SAM3 box segmentation, dept…
暂未收录效果图
查看技能说明shen-shanshan
Generate comprehensive Chinese technical tutorial documents for specific vLLM models (e.g., Qwen3-VL, DeepSeek-V3, Llama 4, InternVL3, etc.). Produces deep-dive model wa…
暂未收录效果图
查看技能说明drpwchen
Lecture recordings → structured grounded notes + a synced HTML viewer: video, timestamped transcript and curated summary on one page. Local GPU pipeline (Whisper ASR · s…