Skill comparison
Use this as a shortlist, then open the skill detail page before adopting.
Decision summary
Strongest overall
ref-hallucination-arena
Shortlist this skill and compare it with close alternatives before production adoption.
Fastest prototype
ref-hallucination-arena
Best first install candidate based on install readiness and adoption.
Freshest repo
ref-hallucination-arena
Most recent maintenance signal among this shortlist.
| Signal | ref-hallucination-arena Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mode. Use when the user asks to evaluate, benchmark, or compare models on academic reference hallucination, literature recommendation quality, or citation accuracy. |
|---|---|
| Quality | 70/100 Strong |
| Decision verdict | 81/100 Strong shortlist Shortlist this skill and compare it with close alternatives before production adoption. |
| Adoption | 809 stars Verified outcomes are shown on each skill page |
| Freshness | Aug 3, 2026 |
| Use-case fit | |
| Workflow fit | |
| Platform hints | Claude Code, OpenAI Agents |
| Warnings | No OpenAgentSkill engagement data yet |
Skill comparison
Use this as a shortlist, then open the skill detail page before adopting.
Decision summary
Strongest overall
ref-hallucination-arena
Shortlist this skill and compare it with close alternatives before production adoption.
Fastest prototype
ref-hallucination-arena
Best first install candidate based on install readiness and adoption.
Freshest repo
ref-hallucination-arena
Most recent maintenance signal among this shortlist.
| Signal | ref-hallucination-arena Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mode. Use when the user asks to evaluate, benchmark, or compare models on academic reference hallucination, literature recommendation quality, or citation accuracy. |
|---|---|
| Quality | 70/100 Strong |
| Decision verdict | 81/100 Strong shortlist Shortlist this skill and compare it with close alternatives before production adoption. |
| Adoption | 809 stars Verified outcomes are shown on each skill page |
| Freshness | Aug 3, 2026 |
| Use-case fit | |
| Workflow fit | |
| Platform hints | Claude Code, OpenAI Agents |
| Warnings | No OpenAgentSkill engagement data yet |
| Best for | Research agents workflows · Claude Code teams · teams that value GitHub adoption signals |
| Not ideal for | teams that need a vendor-supported SLA · high-compliance environments without internal security review |
| OpenAgentSkill engagement | 0 views 0 install copies |
| Install | $ npx skills add agentscope-ai/OpenJudge --skill ref-hallucination-arena |
| Best for | Research agents workflows · Claude Code teams · teams that value GitHub adoption signals |
| Not ideal for | teams that need a vendor-supported SLA · high-compliance environments without internal security review |
| OpenAgentSkill engagement | 0 views 0 install copies |
| Install | $ npx skills add agentscope-ai/OpenJudge --skill ref-hallucination-arena |