Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accu…
$ npx skills add agentscope-ai/OpenJudge --skill ref-hallucination-arenaScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
Claude Code + CLI · 4 targets