Creator · jaechang-hits
Last updated · Sep 3, 2026
Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches.
Creator · jaechang-hits
Last updated · Sep 3, 2026
Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches.
Creator · jaechang-hits
Last updated · Sep 3, 2026
Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches.
Creator · jaechang-hits
Last updated · Sep 3, 2026
Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches.
Sandbox only
Install targets
Codex install prompt
Install the "single-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/legacy/single-cell-annotation. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"jaechang-hits-single-cell-annotation","task":"Install single-cell-annotation","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Maintenance
fresh
8d since push
Risk
Needs review
Permission surface may require sandboxing
GitHub quality
359
72/100 Quality · 78/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
359 GitHub stars
Repo activity
359 stars, 35 forks
Maintenance
8d since push
License
open
Install
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
medium
Skill may inspect schemas, query databases, or work with persistent stores.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
Agent should check
Copy prompt
Task: Use single-cell-annotation in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install
Install command: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
LLM text format
/api/skills/jaechang-hits-single-cell-annotation/install?format=text
Find alternatives
/api/skills/search?q=single-cell-annotation&limit=3
Agent prompt
Use single-cell-annotation for this task. Review https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install, then install with: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/jaechang-hits-single-cell-annotation
LLM text
/api/registry/manifest/jaechang-hits-single-cell-annotation?format=text
Install alias
/api/registry/install/jaechang-hits-single-cell-annotation
Recommend
/api/registry/recommend?task=Use%20single-cell-annotation%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO359 GitHub stars
Stars/forks activity
CHECK359 stars, 35 forks; issue activity unavailable in current metadata
Recent maintenance
PASS8d since push
License clarity
PASSopen
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: single-cell-annotation description: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. license: open ---
# Single Cell RNA-seq Cell Type Annotation
---
## Metadata
**Short Description**: Best practices for annotating cell types in single-cell RNA-seq data using marker-based, automated, and reference-based approaches.
**Authors**: Distilled from "Single-cell best practices" by Luecken, M.D. et al.
**Affiliations**: Helmholtz Munich, Wellcome Sanger Institute, Harvard Medical School, and contributors
**Version**: 1.0
**Last Updated**: January 2025
**License**: CC BY 4.0
**Commercial Use**: ✅ Allowed
**Source**: https://www.sc-best-practices.org/cellular_structure/annotation.html
**Citation**: Luecken, M.D., Theis, F.J. et al. (2023). Current best practices in single-cell RNA-seq analysis: a tutorial. Molecular Systems Biology.
---
## Overview
Cell type annotation is the process of assigning cell type labels to clusters or individual cells in single-cell RNA-seq data. This guide covers three main approaches and their practical implementation.
## Key Concepts
### Cell Type vs. Cell State A **cell type** is a stable identity defined by a developmental trajectory and core marker gene program (e.g., CD4+ T cell, hepatocyte). A **cell state** is a transient condition (activated, cycling, stressed) overlaid on a cell type. Annotation should target cell types first; states are attributes that may further subdivide a type but should not be conflated with type identity.
### Marker Genes and Marker Panels Marker genes are genes whose expression is enriched in a specific cell type relative to other cells in the same tissue context. Reliable annotation uses **panels of multiple markers** (typically 3-5 per type) rather than a single gene, because expression is noisy in droplet-based scRNA-seq and many markers are shared across related types. Markers come in two flavors: **canonical** (literature-derived, e.g., CD3D for T cells) and **data-derived** (from differential expression on the dataset).
### Reference Atlases and Label Transfer A reference atlas is a previously annotated dataset (e.g., Human Cell Atlas, Tabula Sapiens) used to project labels onto a new "query" dataset. Label transfer methods (scArches, scANVI, Azimuth, SingleR) align query cells into the reference latent space and assign the nearest neighbor's label. Quality of transfer depends on tissue match, technology match (e.g., 10x v3 vs. Smart-seq2), and species match.
## Decision Framework
Use this tree to choose an annotation approach:
``` Do you have a well-characterized tissue with a high-quality reference atlas? │ ┌─────────────┴─────────────┐ │ │ YES NO │ │ ▼ ▼ Is this a standard tissue Are you studying (PBMC, lung, gut) with a novel cell types or pre-trained classifier? exploratory data? │ │ ┌─────┴─────┐ ┌─────┴─────┐ │ │ │ │ YES NO YES NO │ │ │ │ ▼ ▼ ▼ ▼ Automated Reference- Manual marker Manual + (CellTypist) based based automated (scArches, (Scanpy, cross-check Azimuth, Seurat) SingleR) ```
### Decision Table
| Scenario | Approach | Primary Tool | Validation | |----------|----------|--------------|------------| | Standard human PBMC, large dataset (>100k cells) | Automated | CellTypist | Spot-check with manual markers | | Well-characterized tissue (lung, kidney, brain) | Reference-based label transfer | scArches / Azimuth | Marker consistency on top clusters | | Novel/rare tissue, no good reference | Manual marker-based | Scanpy / Seurat | Hierarchical, broad-to-fine | | Cross-species (e.g., zebrafish) | Manual markers + ortholog mapping | Scanpy + custom panel | Compare to closest reference species | | Developmental / continuous trajectory | Reference-based with state-aware model | scANVI / scArches | Trajectory coherence + markers | | Disease tissue with known perturbation | Manual + automated cross-check | CellTypist + Scanpy | Confirm disease-specific states separately |
## Three Annotation Approaches
### 1. Manual Marker-Based Annotation Identify cell types by examining expression of known marker genes in each cluster.
**Tools**: Scanpy, Seurat **Best for**: Small datasets, novel cell types, high confidence needs
### 2. Automated Annotation Use pre-trained classifiers to automatically assign cell type labels.
**Tools**: CellTypist, scAnnotate **Best for**: Standard tissues, quick preliminary annotation, large datasets
### 3. Reference-Based Label Transfer Transfer labels from annotated reference datasets to your query data.
**Tools**: scArches, scANVI, Azimuth, SingleR **Best for**: Well-characterized tissues, integration with public data
## Recommended Workflow
### Step 1: Quality Control First - **Remove low-quality cells before annotation** - Filter doublets (expected doublet rate: 0.8% per 1000 cells) - Check for ambient RNA contamination - Verify cluster quality and resolution
### Step 2: Initial Marker-Based Assessment
```python # Scanpy example import scanpy as sc
# Calculate marker genes for clusters sc.tl.rank_genes_groups(adata, 'leiden', method='wilcoxon')
# Visualize top markers sc.pl.rank_genes_groups(adata, n_genes=25, sharey=False)
# Plot known markers markers = { 'T cells': ['CD3D', 'CD3E', 'CD4', 'CD8A'], 'B cells': ['CD19', 'MS4A1', 'CD79A'], 'Monocytes': ['CD14', 'FCGR3A', 'LYZ'], 'NK cells': ['NCAM1', 'NKG7', 'GNLY'] }
sc.pl.dotplot(adata, markers, groupby='leiden') ```
### Step 3: Use Automated Tools for Validation
```python # CellTypist example (fast, accurate for immune cells) import celltypist from celltypist import models
# Download immune cell model model = models.Model.load(model='Immune_All_Low.pkl')
# Predict cell types predictions = celltypist.annotate(adata, model=model, majority_voting=True) adata = predictions.to_adata() ```
### Step 4: Reference-Based Refinement
```python # scArches example for label transfer import scarches as sca
# Load pre-trained reference model model = sca.models.SCANVI.load_query_data( adata=adata, # Your query data reference_model="path/to/reference_model" )
# Transfer labels model.train(max_epochs=100) adata.obs['transferred_labels'] = model.predict() ```
## Best Practices
### Do's: 1. **Always combine multiple approaches** - Use marker-based validation even with automated tools 2. **Check cluster purity** - Ensure clusters represent single cell types 3. **Validate with multiple marker sets** - Don't rely on single markers 4. **Consider biological context** - Tissue type, disease state, developmental stage 5. **Document confidence levels** - Note uncertain annotations 6. **Use hierarchical annotation** - Broad categories first, then subtypes
### Don'ts: 1. **Don't over-cluster** - Too fine resolution creates artificial distinctions 2. **Don't ignore batch effects** - Correct before annotation 3. **Don't trust automation blindly** - Always validate predictions 4. **Don't mix cell states with cell types** - Activated vs. resting cells are states, not types 5. **Don't annotate low-quality cells** - Remove them first
## Common Pitfalls
1. **Doublet Clusters**: Clusters that show markers from multiple cell types are often doublets, not novel hybrid populations. - *How to avoid*: Run doublet detection tools (Scrublet, DoubletFinder) before annotation and remove flagged cells. 2. **Ambient RNA Contamination**: Background markers appear across all cells, blurring cell type boundaries. - *How to avoid*: Apply SoupX or CellBender decontamination during preprocessing — don't trust raw counts on droplet data. 3. **Over-interpretation of Small Clusters**: Rare clusters (<25 cells) are often technical artifacts rather than biological subtypes. - *How to avoid*: Require a minimum cell count threshold and validate with an independent dataset before naming the cluster. 4. **Reference Mismatch**: Transferring labels from a reference built on a different tissue, species, or condition produces confidently wrong annotations. - *How to avoid*: Use tissue- and species-matched references, and check marker-gene overlap between query and reference before label transfer. 5. **Confusing Cell States with Cell Types**: Activated vs. resting T cells, M1 vs. M2 macrophages, and cycling vs. quiescent cells are *states*, not distinct types. - *How to avoid*: Annotate cell type first using stable lineage markers, then layer state annotations on top — don't mix the two axes. 6. **Trusting Automated Tools Blindly**: CellTypist or SingleR predictions look authoritative but can fail silently on out-of-distribution cells. - *How to avoid*: Always cross-check automated calls against marker-based dot plots, and flag low-confidence predictions for manual review. 7. **Annotating Low-Quality Cells**: Including cells with high mitochondrial content or low gene counts contaminates downstream signatures. - *How to avoid*: Apply QC filters (mt%, n_genes, n_counts) before clustering — don't annotate first and clean up later.
## Tool Selection Guide
| Scenario | Recommended Tool | Why | |----------|------------------|-----| | Immune cells (human) | CellTypist | Pre-trained on large immune atlases | | Mouse tissues | scArches + Mouse Cell Atlas | Comprehensive mouse reference | | Novel cell types | Manual + Scanpy/Seurat | Need domain expertise | | Large datasets (>100k cells) | CellTypist | Fast, scalable | | Cross-species | Manual markers | Limited reference transfer | | Developmental data | scArches | Handles continuous states |
## Key Marker Genes by Cell Type
### Blood/Immune: - **T cells**: CD3D, CD3E (all T cells); CD4, CD8A (subtypes) - **B cells**: CD19, MS4A1 (CD20), CD79A - **Monocytes/Macrophages**: CD14, CD68, LYZ - **NK cells**: NCAM1 (CD56), NKG7, KLRD1 - **Dendritic cells**: FCER1A, CD1C
### Epithelial: - **General epithelial**: EPCAM, KRT18, KRT19 - **Lung AT1**: AGER, PDPN - **Lung AT2**: SFTPC, SFTPA1 - **Intestinal**: VIL1, MUC2
### Stromal: - **Fibroblasts**: COL1A1, DCN, LUM - **Endothelial**: PECAM1 (CD31), VWF, CDH5 - **Smooth muscle**: ACTA2, MYH11, TAGLN
## Validation Checklist
- [ ] Cluster purity: >80% cells with same label per cluster - [ ] Marker consistency: Top DE genes match expected markers - [ ] Biological plausibility: Expected proportions for tissue type - [ ] Cross-method agreement: Manual and automated annotations align - [ ] Reference quality: >70% cells successfully transferred - [ ] Doublet check: No clusters with multi-lineage markers - [ ] Documentation: Record confidence levels and uncertain calls
## References
### Tools: - **Scanpy**: https://scanpy.readthedocs.io/ - **CellTypist**: https://www.celltypist.org/ - **scArches**: https://scarches.readthedocs.io/ - **Seurat**: https://satijalab.org/seurat/
### Marker Databases & Atlases: - **PanglaoDB**: https://panglaodb.se/ (Database of marker genes) - **CellMarker**: http://bio-bigdata.hrbmu.edu.cn/CellMarker/ (Curated cell marker database) - **Human Cell Atlas**: https://www.humancellatlas.org/ (Reference datasets) - **Single Cell Best Practices**: https://www.sc-best-practices.org/cellular_structure/annotation.html - **Luecken & Theis (2023)**: Current best practices in single-cell RNA-seq analysis. Molecular Systems Biology.
### Pre-trained Models: - **CellTypist models**: 30+ tissue-specific models - **Azimuth references*
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for single-cell-annotation, ready for a manual X post.
single-cell-annotation: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference... 359 stars https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x
Listing + install path for single-cell-annotation: https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x Install: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to jaechang-hits but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation/audit)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)jaechang-hits
@jaechang-hits
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "single-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/legacy/single-cell-annotation. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"jaechang-hits-single-cell-annotation","task":"Install single-cell-annotation","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Maintenance
fresh
8d since push
Risk
Needs review
Permission surface may require sandboxing
GitHub quality
359
72/100 Quality · 78/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
359 GitHub stars
Repo activity
359 stars, 35 forks
Maintenance
8d since push
License
open
Install
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
medium
Skill may inspect schemas, query databases, or work with persistent stores.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
Agent should check
Copy prompt
Task: Use single-cell-annotation in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install
Install command: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
LLM text format
/api/skills/jaechang-hits-single-cell-annotation/install?format=text
Find alternatives
/api/skills/search?q=single-cell-annotation&limit=3
Agent prompt
Use single-cell-annotation for this task. Review https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install, then install with: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/jaechang-hits-single-cell-annotation
LLM text
/api/registry/manifest/jaechang-hits-single-cell-annotation?format=text
Install alias
/api/registry/install/jaechang-hits-single-cell-annotation
Recommend
/api/registry/recommend?task=Use%20single-cell-annotation%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO359 GitHub stars
Stars/forks activity
CHECK359 stars, 35 forks; issue activity unavailable in current metadata
Recent maintenance
PASS8d since push
License clarity
PASSopen
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: single-cell-annotation description: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. license: open ---
# Single Cell RNA-seq Cell Type Annotation
---
## Metadata
**Short Description**: Best practices for annotating cell types in single-cell RNA-seq data using marker-based, automated, and reference-based approaches.
**Authors**: Distilled from "Single-cell best practices" by Luecken, M.D. et al.
**Affiliations**: Helmholtz Munich, Wellcome Sanger Institute, Harvard Medical School, and contributors
**Version**: 1.0
**Last Updated**: January 2025
**License**: CC BY 4.0
**Commercial Use**: ✅ Allowed
**Source**: https://www.sc-best-practices.org/cellular_structure/annotation.html
**Citation**: Luecken, M.D., Theis, F.J. et al. (2023). Current best practices in single-cell RNA-seq analysis: a tutorial. Molecular Systems Biology.
---
## Overview
Cell type annotation is the process of assigning cell type labels to clusters or individual cells in single-cell RNA-seq data. This guide covers three main approaches and their practical implementation.
## Key Concepts
### Cell Type vs. Cell State A **cell type** is a stable identity defined by a developmental trajectory and core marker gene program (e.g., CD4+ T cell, hepatocyte). A **cell state** is a transient condition (activated, cycling, stressed) overlaid on a cell type. Annotation should target cell types first; states are attributes that may further subdivide a type but should not be conflated with type identity.
### Marker Genes and Marker Panels Marker genes are genes whose expression is enriched in a specific cell type relative to other cells in the same tissue context. Reliable annotation uses **panels of multiple markers** (typically 3-5 per type) rather than a single gene, because expression is noisy in droplet-based scRNA-seq and many markers are shared across related types. Markers come in two flavors: **canonical** (literature-derived, e.g., CD3D for T cells) and **data-derived** (from differential expression on the dataset).
### Reference Atlases and Label Transfer A reference atlas is a previously annotated dataset (e.g., Human Cell Atlas, Tabula Sapiens) used to project labels onto a new "query" dataset. Label transfer methods (scArches, scANVI, Azimuth, SingleR) align query cells into the reference latent space and assign the nearest neighbor's label. Quality of transfer depends on tissue match, technology match (e.g., 10x v3 vs. Smart-seq2), and species match.
## Decision Framework
Use this tree to choose an annotation approach:
``` Do you have a well-characterized tissue with a high-quality reference atlas? │ ┌─────────────┴─────────────┐ │ │ YES NO │ │ ▼ ▼ Is this a standard tissue Are you studying (PBMC, lung, gut) with a novel cell types or pre-trained classifier? exploratory data? │ │ ┌─────┴─────┐ ┌─────┴─────┐ │ │ │ │ YES NO YES NO │ │ │ │ ▼ ▼ ▼ ▼ Automated Reference- Manual marker Manual + (CellTypist) based based automated (scArches, (Scanpy, cross-check Azimuth, Seurat) SingleR) ```
### Decision Table
| Scenario | Approach | Primary Tool | Validation | |----------|----------|--------------|------------| | Standard human PBMC, large dataset (>100k cells) | Automated | CellTypist | Spot-check with manual markers | | Well-characterized tissue (lung, kidney, brain) | Reference-based label transfer | scArches / Azimuth | Marker consistency on top clusters | | Novel/rare tissue, no good reference | Manual marker-based | Scanpy / Seurat | Hierarchical, broad-to-fine | | Cross-species (e.g., zebrafish) | Manual markers + ortholog mapping | Scanpy + custom panel | Compare to closest reference species | | Developmental / continuous trajectory | Reference-based with state-aware model | scANVI / scArches | Trajectory coherence + markers | | Disease tissue with known perturbation | Manual + automated cross-check | CellTypist + Scanpy | Confirm disease-specific states separately |
## Three Annotation Approaches
### 1. Manual Marker-Based Annotation Identify cell types by examining expression of known marker genes in each cluster.
**Tools**: Scanpy, Seurat **Best for**: Small datasets, novel cell types, high confidence needs
### 2. Automated Annotation Use pre-trained classifiers to automatically assign cell type labels.
**Tools**: CellTypist, scAnnotate **Best for**: Standard tissues, quick preliminary annotation, large datasets
### 3. Reference-Based Label Transfer Transfer labels from annotated reference datasets to your query data.
**Tools**: scArches, scANVI, Azimuth, SingleR **Best for**: Well-characterized tissues, integration with public data
## Recommended Workflow
### Step 1: Quality Control First - **Remove low-quality cells before annotation** - Filter doublets (expected doublet rate: 0.8% per 1000 cells) - Check for ambient RNA contamination - Verify cluster quality and resolution
### Step 2: Initial Marker-Based Assessment
```python # Scanpy example import scanpy as sc
# Calculate marker genes for clusters sc.tl.rank_genes_groups(adata, 'leiden', method='wilcoxon')
# Visualize top markers sc.pl.rank_genes_groups(adata, n_genes=25, sharey=False)
# Plot known markers markers = { 'T cells': ['CD3D', 'CD3E', 'CD4', 'CD8A'], 'B cells': ['CD19', 'MS4A1', 'CD79A'], 'Monocytes': ['CD14', 'FCGR3A', 'LYZ'], 'NK cells': ['NCAM1', 'NKG7', 'GNLY'] }
sc.pl.dotplot(adata, markers, groupby='leiden') ```
### Step 3: Use Automated Tools for Validation
```python # CellTypist example (fast, accurate for immune cells) import celltypist from celltypist import models
# Download immune cell model model = models.Model.load(model='Immune_All_Low.pkl')
# Predict cell types predictions = celltypist.annotate(adata, model=model, majority_voting=True) adata = predictions.to_adata() ```
### Step 4: Reference-Based Refinement
```python # scArches example for label transfer import scarches as sca
# Load pre-trained reference model model = sca.models.SCANVI.load_query_data( adata=adata, # Your query data reference_model="path/to/reference_model" )
# Transfer labels model.train(max_epochs=100) adata.obs['transferred_labels'] = model.predict() ```
## Best Practices
### Do's: 1. **Always combine multiple approaches** - Use marker-based validation even with automated tools 2. **Check cluster purity** - Ensure clusters represent single cell types 3. **Validate with multiple marker sets** - Don't rely on single markers 4. **Consider biological context** - Tissue type, disease state, developmental stage 5. **Document confidence levels** - Note uncertain annotations 6. **Use hierarchical annotation** - Broad categories first, then subtypes
### Don'ts: 1. **Don't over-cluster** - Too fine resolution creates artificial distinctions 2. **Don't ignore batch effects** - Correct before annotation 3. **Don't trust automation blindly** - Always validate predictions 4. **Don't mix cell states with cell types** - Activated vs. resting cells are states, not types 5. **Don't annotate low-quality cells** - Remove them first
## Common Pitfalls
1. **Doublet Clusters**: Clusters that show markers from multiple cell types are often doublets, not novel hybrid populations. - *How to avoid*: Run doublet detection tools (Scrublet, DoubletFinder) before annotation and remove flagged cells. 2. **Ambient RNA Contamination**: Background markers appear across all cells, blurring cell type boundaries. - *How to avoid*: Apply SoupX or CellBender decontamination during preprocessing — don't trust raw counts on droplet data. 3. **Over-interpretation of Small Clusters**: Rare clusters (<25 cells) are often technical artifacts rather than biological subtypes. - *How to avoid*: Require a minimum cell count threshold and validate with an independent dataset before naming the cluster. 4. **Reference Mismatch**: Transferring labels from a reference built on a different tissue, species, or condition produces confidently wrong annotations. - *How to avoid*: Use tissue- and species-matched references, and check marker-gene overlap between query and reference before label transfer. 5. **Confusing Cell States with Cell Types**: Activated vs. resting T cells, M1 vs. M2 macrophages, and cycling vs. quiescent cells are *states*, not distinct types. - *How to avoid*: Annotate cell type first using stable lineage markers, then layer state annotations on top — don't mix the two axes. 6. **Trusting Automated Tools Blindly**: CellTypist or SingleR predictions look authoritative but can fail silently on out-of-distribution cells. - *How to avoid*: Always cross-check automated calls against marker-based dot plots, and flag low-confidence predictions for manual review. 7. **Annotating Low-Quality Cells**: Including cells with high mitochondrial content or low gene counts contaminates downstream signatures. - *How to avoid*: Apply QC filters (mt%, n_genes, n_counts) before clustering — don't annotate first and clean up later.
## Tool Selection Guide
| Scenario | Recommended Tool | Why | |----------|------------------|-----| | Immune cells (human) | CellTypist | Pre-trained on large immune atlases | | Mouse tissues | scArches + Mouse Cell Atlas | Comprehensive mouse reference | | Novel cell types | Manual + Scanpy/Seurat | Need domain expertise | | Large datasets (>100k cells) | CellTypist | Fast, scalable | | Cross-species | Manual markers | Limited reference transfer | | Developmental data | scArches | Handles continuous states |
## Key Marker Genes by Cell Type
### Blood/Immune: - **T cells**: CD3D, CD3E (all T cells); CD4, CD8A (subtypes) - **B cells**: CD19, MS4A1 (CD20), CD79A - **Monocytes/Macrophages**: CD14, CD68, LYZ - **NK cells**: NCAM1 (CD56), NKG7, KLRD1 - **Dendritic cells**: FCER1A, CD1C
### Epithelial: - **General epithelial**: EPCAM, KRT18, KRT19 - **Lung AT1**: AGER, PDPN - **Lung AT2**: SFTPC, SFTPA1 - **Intestinal**: VIL1, MUC2
### Stromal: - **Fibroblasts**: COL1A1, DCN, LUM - **Endothelial**: PECAM1 (CD31), VWF, CDH5 - **Smooth muscle**: ACTA2, MYH11, TAGLN
## Validation Checklist
- [ ] Cluster purity: >80% cells with same label per cluster - [ ] Marker consistency: Top DE genes match expected markers - [ ] Biological plausibility: Expected proportions for tissue type - [ ] Cross-method agreement: Manual and automated annotations align - [ ] Reference quality: >70% cells successfully transferred - [ ] Doublet check: No clusters with multi-lineage markers - [ ] Documentation: Record confidence levels and uncertain calls
## References
### Tools: - **Scanpy**: https://scanpy.readthedocs.io/ - **CellTypist**: https://www.celltypist.org/ - **scArches**: https://scarches.readthedocs.io/ - **Seurat**: https://satijalab.org/seurat/
### Marker Databases & Atlases: - **PanglaoDB**: https://panglaodb.se/ (Database of marker genes) - **CellMarker**: http://bio-bigdata.hrbmu.edu.cn/CellMarker/ (Curated cell marker database) - **Human Cell Atlas**: https://www.humancellatlas.org/ (Reference datasets) - **Single Cell Best Practices**: https://www.sc-best-practices.org/cellular_structure/annotation.html - **Luecken & Theis (2023)**: Current best practices in single-cell RNA-seq analysis. Molecular Systems Biology.
### Pre-trained Models: - **CellTypist models**: 30+ tissue-specific models - **Azimuth references*
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for single-cell-annotation, ready for a manual X post.
single-cell-annotation: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference... 359 stars https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x
Listing + install path for single-cell-annotation: https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x Install: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to jaechang-hits but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation/audit)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)jaechang-hits
@jaechang-hits
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "single-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/legacy/single-cell-annotation. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"jaechang-hits-single-cell-annotation","task":"Install single-cell-annotation","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Maintenance
fresh
8d since push
Risk
Needs review
Permission surface may require sandboxing
GitHub quality
359
72/100 Quality · 78/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
359 GitHub stars
Repo activity
359 stars, 35 forks
Maintenance
8d since push
License
open
Install
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
medium
Skill may inspect schemas, query databases, or work with persistent stores.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
Agent should check
Copy prompt
Task: Use single-cell-annotation in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install
Install command: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
LLM text format
/api/skills/jaechang-hits-single-cell-annotation/install?format=text
Find alternatives
/api/skills/search?q=single-cell-annotation&limit=3
Agent prompt
Use single-cell-annotation for this task. Review https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install, then install with: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/jaechang-hits-single-cell-annotation
LLM text
/api/registry/manifest/jaechang-hits-single-cell-annotation?format=text
Install alias
/api/registry/install/jaechang-hits-single-cell-annotation
Recommend
/api/registry/recommend?task=Use%20single-cell-annotation%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO359 GitHub stars
Stars/forks activity
CHECK359 stars, 35 forks; issue activity unavailable in current metadata
Recent maintenance
PASS8d since push
License clarity
PASSopen
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: single-cell-annotation description: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. license: open ---
# Single Cell RNA-seq Cell Type Annotation
---
## Metadata
**Short Description**: Best practices for annotating cell types in single-cell RNA-seq data using marker-based, automated, and reference-based approaches.
**Authors**: Distilled from "Single-cell best practices" by Luecken, M.D. et al.
**Affiliations**: Helmholtz Munich, Wellcome Sanger Institute, Harvard Medical School, and contributors
**Version**: 1.0
**Last Updated**: January 2025
**License**: CC BY 4.0
**Commercial Use**: ✅ Allowed
**Source**: https://www.sc-best-practices.org/cellular_structure/annotation.html
**Citation**: Luecken, M.D., Theis, F.J. et al. (2023). Current best practices in single-cell RNA-seq analysis: a tutorial. Molecular Systems Biology.
---
## Overview
Cell type annotation is the process of assigning cell type labels to clusters or individual cells in single-cell RNA-seq data. This guide covers three main approaches and their practical implementation.
## Key Concepts
### Cell Type vs. Cell State A **cell type** is a stable identity defined by a developmental trajectory and core marker gene program (e.g., CD4+ T cell, hepatocyte). A **cell state** is a transient condition (activated, cycling, stressed) overlaid on a cell type. Annotation should target cell types first; states are attributes that may further subdivide a type but should not be conflated with type identity.
### Marker Genes and Marker Panels Marker genes are genes whose expression is enriched in a specific cell type relative to other cells in the same tissue context. Reliable annotation uses **panels of multiple markers** (typically 3-5 per type) rather than a single gene, because expression is noisy in droplet-based scRNA-seq and many markers are shared across related types. Markers come in two flavors: **canonical** (literature-derived, e.g., CD3D for T cells) and **data-derived** (from differential expression on the dataset).
### Reference Atlases and Label Transfer A reference atlas is a previously annotated dataset (e.g., Human Cell Atlas, Tabula Sapiens) used to project labels onto a new "query" dataset. Label transfer methods (scArches, scANVI, Azimuth, SingleR) align query cells into the reference latent space and assign the nearest neighbor's label. Quality of transfer depends on tissue match, technology match (e.g., 10x v3 vs. Smart-seq2), and species match.
## Decision Framework
Use this tree to choose an annotation approach:
``` Do you have a well-characterized tissue with a high-quality reference atlas? │ ┌─────────────┴─────────────┐ │ │ YES NO │ │ ▼ ▼ Is this a standard tissue Are you studying (PBMC, lung, gut) with a novel cell types or pre-trained classifier? exploratory data? │ │ ┌─────┴─────┐ ┌─────┴─────┐ │ │ │ │ YES NO YES NO │ │ │ │ ▼ ▼ ▼ ▼ Automated Reference- Manual marker Manual + (CellTypist) based based automated (scArches, (Scanpy, cross-check Azimuth, Seurat) SingleR) ```
### Decision Table
| Scenario | Approach | Primary Tool | Validation | |----------|----------|--------------|------------| | Standard human PBMC, large dataset (>100k cells) | Automated | CellTypist | Spot-check with manual markers | | Well-characterized tissue (lung, kidney, brain) | Reference-based label transfer | scArches / Azimuth | Marker consistency on top clusters | | Novel/rare tissue, no good reference | Manual marker-based | Scanpy / Seurat | Hierarchical, broad-to-fine | | Cross-species (e.g., zebrafish) | Manual markers + ortholog mapping | Scanpy + custom panel | Compare to closest reference species | | Developmental / continuous trajectory | Reference-based with state-aware model | scANVI / scArches | Trajectory coherence + markers | | Disease tissue with known perturbation | Manual + automated cross-check | CellTypist + Scanpy | Confirm disease-specific states separately |
## Three Annotation Approaches
### 1. Manual Marker-Based Annotation Identify cell types by examining expression of known marker genes in each cluster.
**Tools**: Scanpy, Seurat **Best for**: Small datasets, novel cell types, high confidence needs
### 2. Automated Annotation Use pre-trained classifiers to automatically assign cell type labels.
**Tools**: CellTypist, scAnnotate **Best for**: Standard tissues, quick preliminary annotation, large datasets
### 3. Reference-Based Label Transfer Transfer labels from annotated reference datasets to your query data.
**Tools**: scArches, scANVI, Azimuth, SingleR **Best for**: Well-characterized tissues, integration with public data
## Recommended Workflow
### Step 1: Quality Control First - **Remove low-quality cells before annotation** - Filter doublets (expected doublet rate: 0.8% per 1000 cells) - Check for ambient RNA contamination - Verify cluster quality and resolution
### Step 2: Initial Marker-Based Assessment
```python # Scanpy example import scanpy as sc
# Calculate marker genes for clusters sc.tl.rank_genes_groups(adata, 'leiden', method='wilcoxon')
# Visualize top markers sc.pl.rank_genes_groups(adata, n_genes=25, sharey=False)
# Plot known markers markers = { 'T cells': ['CD3D', 'CD3E', 'CD4', 'CD8A'], 'B cells': ['CD19', 'MS4A1', 'CD79A'], 'Monocytes': ['CD14', 'FCGR3A', 'LYZ'], 'NK cells': ['NCAM1', 'NKG7', 'GNLY'] }
sc.pl.dotplot(adata, markers, groupby='leiden') ```
### Step 3: Use Automated Tools for Validation
```python # CellTypist example (fast, accurate for immune cells) import celltypist from celltypist import models
# Download immune cell model model = models.Model.load(model='Immune_All_Low.pkl')
# Predict cell types predictions = celltypist.annotate(adata, model=model, majority_voting=True) adata = predictions.to_adata() ```
### Step 4: Reference-Based Refinement
```python # scArches example for label transfer import scarches as sca
# Load pre-trained reference model model = sca.models.SCANVI.load_query_data( adata=adata, # Your query data reference_model="path/to/reference_model" )
# Transfer labels model.train(max_epochs=100) adata.obs['transferred_labels'] = model.predict() ```
## Best Practices
### Do's: 1. **Always combine multiple approaches** - Use marker-based validation even with automated tools 2. **Check cluster purity** - Ensure clusters represent single cell types 3. **Validate with multiple marker sets** - Don't rely on single markers 4. **Consider biological context** - Tissue type, disease state, developmental stage 5. **Document confidence levels** - Note uncertain annotations 6. **Use hierarchical annotation** - Broad categories first, then subtypes
### Don'ts: 1. **Don't over-cluster** - Too fine resolution creates artificial distinctions 2. **Don't ignore batch effects** - Correct before annotation 3. **Don't trust automation blindly** - Always validate predictions 4. **Don't mix cell states with cell types** - Activated vs. resting cells are states, not types 5. **Don't annotate low-quality cells** - Remove them first
## Common Pitfalls
1. **Doublet Clusters**: Clusters that show markers from multiple cell types are often doublets, not novel hybrid populations. - *How to avoid*: Run doublet detection tools (Scrublet, DoubletFinder) before annotation and remove flagged cells. 2. **Ambient RNA Contamination**: Background markers appear across all cells, blurring cell type boundaries. - *How to avoid*: Apply SoupX or CellBender decontamination during preprocessing — don't trust raw counts on droplet data. 3. **Over-interpretation of Small Clusters**: Rare clusters (<25 cells) are often technical artifacts rather than biological subtypes. - *How to avoid*: Require a minimum cell count threshold and validate with an independent dataset before naming the cluster. 4. **Reference Mismatch**: Transferring labels from a reference built on a different tissue, species, or condition produces confidently wrong annotations. - *How to avoid*: Use tissue- and species-matched references, and check marker-gene overlap between query and reference before label transfer. 5. **Confusing Cell States with Cell Types**: Activated vs. resting T cells, M1 vs. M2 macrophages, and cycling vs. quiescent cells are *states*, not distinct types. - *How to avoid*: Annotate cell type first using stable lineage markers, then layer state annotations on top — don't mix the two axes. 6. **Trusting Automated Tools Blindly**: CellTypist or SingleR predictions look authoritative but can fail silently on out-of-distribution cells. - *How to avoid*: Always cross-check automated calls against marker-based dot plots, and flag low-confidence predictions for manual review. 7. **Annotating Low-Quality Cells**: Including cells with high mitochondrial content or low gene counts contaminates downstream signatures. - *How to avoid*: Apply QC filters (mt%, n_genes, n_counts) before clustering — don't annotate first and clean up later.
## Tool Selection Guide
| Scenario | Recommended Tool | Why | |----------|------------------|-----| | Immune cells (human) | CellTypist | Pre-trained on large immune atlases | | Mouse tissues | scArches + Mouse Cell Atlas | Comprehensive mouse reference | | Novel cell types | Manual + Scanpy/Seurat | Need domain expertise | | Large datasets (>100k cells) | CellTypist | Fast, scalable | | Cross-species | Manual markers | Limited reference transfer | | Developmental data | scArches | Handles continuous states |
## Key Marker Genes by Cell Type
### Blood/Immune: - **T cells**: CD3D, CD3E (all T cells); CD4, CD8A (subtypes) - **B cells**: CD19, MS4A1 (CD20), CD79A - **Monocytes/Macrophages**: CD14, CD68, LYZ - **NK cells**: NCAM1 (CD56), NKG7, KLRD1 - **Dendritic cells**: FCER1A, CD1C
### Epithelial: - **General epithelial**: EPCAM, KRT18, KRT19 - **Lung AT1**: AGER, PDPN - **Lung AT2**: SFTPC, SFTPA1 - **Intestinal**: VIL1, MUC2
### Stromal: - **Fibroblasts**: COL1A1, DCN, LUM - **Endothelial**: PECAM1 (CD31), VWF, CDH5 - **Smooth muscle**: ACTA2, MYH11, TAGLN
## Validation Checklist
- [ ] Cluster purity: >80% cells with same label per cluster - [ ] Marker consistency: Top DE genes match expected markers - [ ] Biological plausibility: Expected proportions for tissue type - [ ] Cross-method agreement: Manual and automated annotations align - [ ] Reference quality: >70% cells successfully transferred - [ ] Doublet check: No clusters with multi-lineage markers - [ ] Documentation: Record confidence levels and uncertain calls
## References
### Tools: - **Scanpy**: https://scanpy.readthedocs.io/ - **CellTypist**: https://www.celltypist.org/ - **scArches**: https://scarches.readthedocs.io/ - **Seurat**: https://satijalab.org/seurat/
### Marker Databases & Atlases: - **PanglaoDB**: https://panglaodb.se/ (Database of marker genes) - **CellMarker**: http://bio-bigdata.hrbmu.edu.cn/CellMarker/ (Curated cell marker database) - **Human Cell Atlas**: https://www.humancellatlas.org/ (Reference datasets) - **Single Cell Best Practices**: https://www.sc-best-practices.org/cellular_structure/annotation.html - **Luecken & Theis (2023)**: Current best practices in single-cell RNA-seq analysis. Molecular Systems Biology.
### Pre-trained Models: - **CellTypist models**: 30+ tissue-specific models - **Azimuth references*
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for single-cell-annotation, ready for a manual X post.
single-cell-annotation: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference... 359 stars https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x
Listing + install path for single-cell-annotation: https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x Install: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to jaechang-hits but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation/audit)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)jaechang-hits
@jaechang-hits
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "single-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/legacy/single-cell-annotation. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"jaechang-hits-single-cell-annotation","task":"Install single-cell-annotation","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Maintenance
fresh
8d since push
Risk
Needs review
Permission surface may require sandboxing
GitHub quality
359
72/100 Quality · 78/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
359 GitHub stars
Repo activity
359 stars, 35 forks
Maintenance
8d since push
License
open
Install
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
medium
Skill may inspect schemas, query databases, or work with persistent stores.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
Agent should check
Copy prompt
Task: Use single-cell-annotation in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20single-cell-annotation%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install
Install command: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/jaechang-hits-single-cell-annotation/install
LLM text format
/api/skills/jaechang-hits-single-cell-annotation/install?format=text
Find alternatives
/api/skills/search?q=single-cell-annotation&limit=3
Agent prompt
Use single-cell-annotation for this task. Review https://www.openagentskill.com/api/skills/jaechang-hits-single-cell-annotation/install, then install with: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotationRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/jaechang-hits-single-cell-annotation
LLM text
/api/registry/manifest/jaechang-hits-single-cell-annotation?format=text
Install alias
/api/registry/install/jaechang-hits-single-cell-annotation
Recommend
/api/registry/recommend?task=Use%20single-cell-annotation%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO359 GitHub stars
Stars/forks activity
CHECK359 stars, 35 forks; issue activity unavailable in current metadata
Recent maintenance
PASS8d since push
License clarity
PASSopen
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: single-cell-annotation description: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference-based, and automated classification approaches. license: open ---
# Single Cell RNA-seq Cell Type Annotation
---
## Metadata
**Short Description**: Best practices for annotating cell types in single-cell RNA-seq data using marker-based, automated, and reference-based approaches.
**Authors**: Distilled from "Single-cell best practices" by Luecken, M.D. et al.
**Affiliations**: Helmholtz Munich, Wellcome Sanger Institute, Harvard Medical School, and contributors
**Version**: 1.0
**Last Updated**: January 2025
**License**: CC BY 4.0
**Commercial Use**: ✅ Allowed
**Source**: https://www.sc-best-practices.org/cellular_structure/annotation.html
**Citation**: Luecken, M.D., Theis, F.J. et al. (2023). Current best practices in single-cell RNA-seq analysis: a tutorial. Molecular Systems Biology.
---
## Overview
Cell type annotation is the process of assigning cell type labels to clusters or individual cells in single-cell RNA-seq data. This guide covers three main approaches and their practical implementation.
## Key Concepts
### Cell Type vs. Cell State A **cell type** is a stable identity defined by a developmental trajectory and core marker gene program (e.g., CD4+ T cell, hepatocyte). A **cell state** is a transient condition (activated, cycling, stressed) overlaid on a cell type. Annotation should target cell types first; states are attributes that may further subdivide a type but should not be conflated with type identity.
### Marker Genes and Marker Panels Marker genes are genes whose expression is enriched in a specific cell type relative to other cells in the same tissue context. Reliable annotation uses **panels of multiple markers** (typically 3-5 per type) rather than a single gene, because expression is noisy in droplet-based scRNA-seq and many markers are shared across related types. Markers come in two flavors: **canonical** (literature-derived, e.g., CD3D for T cells) and **data-derived** (from differential expression on the dataset).
### Reference Atlases and Label Transfer A reference atlas is a previously annotated dataset (e.g., Human Cell Atlas, Tabula Sapiens) used to project labels onto a new "query" dataset. Label transfer methods (scArches, scANVI, Azimuth, SingleR) align query cells into the reference latent space and assign the nearest neighbor's label. Quality of transfer depends on tissue match, technology match (e.g., 10x v3 vs. Smart-seq2), and species match.
## Decision Framework
Use this tree to choose an annotation approach:
``` Do you have a well-characterized tissue with a high-quality reference atlas? │ ┌─────────────┴─────────────┐ │ │ YES NO │ │ ▼ ▼ Is this a standard tissue Are you studying (PBMC, lung, gut) with a novel cell types or pre-trained classifier? exploratory data? │ │ ┌─────┴─────┐ ┌─────┴─────┐ │ │ │ │ YES NO YES NO │ │ │ │ ▼ ▼ ▼ ▼ Automated Reference- Manual marker Manual + (CellTypist) based based automated (scArches, (Scanpy, cross-check Azimuth, Seurat) SingleR) ```
### Decision Table
| Scenario | Approach | Primary Tool | Validation | |----------|----------|--------------|------------| | Standard human PBMC, large dataset (>100k cells) | Automated | CellTypist | Spot-check with manual markers | | Well-characterized tissue (lung, kidney, brain) | Reference-based label transfer | scArches / Azimuth | Marker consistency on top clusters | | Novel/rare tissue, no good reference | Manual marker-based | Scanpy / Seurat | Hierarchical, broad-to-fine | | Cross-species (e.g., zebrafish) | Manual markers + ortholog mapping | Scanpy + custom panel | Compare to closest reference species | | Developmental / continuous trajectory | Reference-based with state-aware model | scANVI / scArches | Trajectory coherence + markers | | Disease tissue with known perturbation | Manual + automated cross-check | CellTypist + Scanpy | Confirm disease-specific states separately |
## Three Annotation Approaches
### 1. Manual Marker-Based Annotation Identify cell types by examining expression of known marker genes in each cluster.
**Tools**: Scanpy, Seurat **Best for**: Small datasets, novel cell types, high confidence needs
### 2. Automated Annotation Use pre-trained classifiers to automatically assign cell type labels.
**Tools**: CellTypist, scAnnotate **Best for**: Standard tissues, quick preliminary annotation, large datasets
### 3. Reference-Based Label Transfer Transfer labels from annotated reference datasets to your query data.
**Tools**: scArches, scANVI, Azimuth, SingleR **Best for**: Well-characterized tissues, integration with public data
## Recommended Workflow
### Step 1: Quality Control First - **Remove low-quality cells before annotation** - Filter doublets (expected doublet rate: 0.8% per 1000 cells) - Check for ambient RNA contamination - Verify cluster quality and resolution
### Step 2: Initial Marker-Based Assessment
```python # Scanpy example import scanpy as sc
# Calculate marker genes for clusters sc.tl.rank_genes_groups(adata, 'leiden', method='wilcoxon')
# Visualize top markers sc.pl.rank_genes_groups(adata, n_genes=25, sharey=False)
# Plot known markers markers = { 'T cells': ['CD3D', 'CD3E', 'CD4', 'CD8A'], 'B cells': ['CD19', 'MS4A1', 'CD79A'], 'Monocytes': ['CD14', 'FCGR3A', 'LYZ'], 'NK cells': ['NCAM1', 'NKG7', 'GNLY'] }
sc.pl.dotplot(adata, markers, groupby='leiden') ```
### Step 3: Use Automated Tools for Validation
```python # CellTypist example (fast, accurate for immune cells) import celltypist from celltypist import models
# Download immune cell model model = models.Model.load(model='Immune_All_Low.pkl')
# Predict cell types predictions = celltypist.annotate(adata, model=model, majority_voting=True) adata = predictions.to_adata() ```
### Step 4: Reference-Based Refinement
```python # scArches example for label transfer import scarches as sca
# Load pre-trained reference model model = sca.models.SCANVI.load_query_data( adata=adata, # Your query data reference_model="path/to/reference_model" )
# Transfer labels model.train(max_epochs=100) adata.obs['transferred_labels'] = model.predict() ```
## Best Practices
### Do's: 1. **Always combine multiple approaches** - Use marker-based validation even with automated tools 2. **Check cluster purity** - Ensure clusters represent single cell types 3. **Validate with multiple marker sets** - Don't rely on single markers 4. **Consider biological context** - Tissue type, disease state, developmental stage 5. **Document confidence levels** - Note uncertain annotations 6. **Use hierarchical annotation** - Broad categories first, then subtypes
### Don'ts: 1. **Don't over-cluster** - Too fine resolution creates artificial distinctions 2. **Don't ignore batch effects** - Correct before annotation 3. **Don't trust automation blindly** - Always validate predictions 4. **Don't mix cell states with cell types** - Activated vs. resting cells are states, not types 5. **Don't annotate low-quality cells** - Remove them first
## Common Pitfalls
1. **Doublet Clusters**: Clusters that show markers from multiple cell types are often doublets, not novel hybrid populations. - *How to avoid*: Run doublet detection tools (Scrublet, DoubletFinder) before annotation and remove flagged cells. 2. **Ambient RNA Contamination**: Background markers appear across all cells, blurring cell type boundaries. - *How to avoid*: Apply SoupX or CellBender decontamination during preprocessing — don't trust raw counts on droplet data. 3. **Over-interpretation of Small Clusters**: Rare clusters (<25 cells) are often technical artifacts rather than biological subtypes. - *How to avoid*: Require a minimum cell count threshold and validate with an independent dataset before naming the cluster. 4. **Reference Mismatch**: Transferring labels from a reference built on a different tissue, species, or condition produces confidently wrong annotations. - *How to avoid*: Use tissue- and species-matched references, and check marker-gene overlap between query and reference before label transfer. 5. **Confusing Cell States with Cell Types**: Activated vs. resting T cells, M1 vs. M2 macrophages, and cycling vs. quiescent cells are *states*, not distinct types. - *How to avoid*: Annotate cell type first using stable lineage markers, then layer state annotations on top — don't mix the two axes. 6. **Trusting Automated Tools Blindly**: CellTypist or SingleR predictions look authoritative but can fail silently on out-of-distribution cells. - *How to avoid*: Always cross-check automated calls against marker-based dot plots, and flag low-confidence predictions for manual review. 7. **Annotating Low-Quality Cells**: Including cells with high mitochondrial content or low gene counts contaminates downstream signatures. - *How to avoid*: Apply QC filters (mt%, n_genes, n_counts) before clustering — don't annotate first and clean up later.
## Tool Selection Guide
| Scenario | Recommended Tool | Why | |----------|------------------|-----| | Immune cells (human) | CellTypist | Pre-trained on large immune atlases | | Mouse tissues | scArches + Mouse Cell Atlas | Comprehensive mouse reference | | Novel cell types | Manual + Scanpy/Seurat | Need domain expertise | | Large datasets (>100k cells) | CellTypist | Fast, scalable | | Cross-species | Manual markers | Limited reference transfer | | Developmental data | scArches | Handles continuous states |
## Key Marker Genes by Cell Type
### Blood/Immune: - **T cells**: CD3D, CD3E (all T cells); CD4, CD8A (subtypes) - **B cells**: CD19, MS4A1 (CD20), CD79A - **Monocytes/Macrophages**: CD14, CD68, LYZ - **NK cells**: NCAM1 (CD56), NKG7, KLRD1 - **Dendritic cells**: FCER1A, CD1C
### Epithelial: - **General epithelial**: EPCAM, KRT18, KRT19 - **Lung AT1**: AGER, PDPN - **Lung AT2**: SFTPC, SFTPA1 - **Intestinal**: VIL1, MUC2
### Stromal: - **Fibroblasts**: COL1A1, DCN, LUM - **Endothelial**: PECAM1 (CD31), VWF, CDH5 - **Smooth muscle**: ACTA2, MYH11, TAGLN
## Validation Checklist
- [ ] Cluster purity: >80% cells with same label per cluster - [ ] Marker consistency: Top DE genes match expected markers - [ ] Biological plausibility: Expected proportions for tissue type - [ ] Cross-method agreement: Manual and automated annotations align - [ ] Reference quality: >70% cells successfully transferred - [ ] Doublet check: No clusters with multi-lineage markers - [ ] Documentation: Record confidence levels and uncertain calls
## References
### Tools: - **Scanpy**: https://scanpy.readthedocs.io/ - **CellTypist**: https://www.celltypist.org/ - **scArches**: https://scarches.readthedocs.io/ - **Seurat**: https://satijalab.org/seurat/
### Marker Databases & Atlases: - **PanglaoDB**: https://panglaodb.se/ (Database of marker genes) - **CellMarker**: http://bio-bigdata.hrbmu.edu.cn/CellMarker/ (Curated cell marker database) - **Human Cell Atlas**: https://www.humancellatlas.org/ (Reference datasets) - **Single Cell Best Practices**: https://www.sc-best-practices.org/cellular_structure/annotation.html - **Luecken & Theis (2023)**: Current best practices in single-cell RNA-seq analysis. Molecular Systems Biology.
### Pre-trained Models: - **CellTypist models**: 30+ tissue-specific models - **Azimuth references*
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for single-cell-annotation, ready for a manual X post.
single-cell-annotation: Best practices for single-cell RNA-seq cell type annotation including marker-based, reference... 359 stars https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x
Listing + install path for single-cell-annotation: https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=x Install: npx skills add jaechang-hits/SciAgent-Skills --skill single-cell-annotation
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to jaechang-hits but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation/audit)
[](https://www.openagentskill.com/skills/jaechang-hits-single-cell-annotation?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)jaechang-hits
@jaechang-hits
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsPermission surface
filesystem or document access, network or browser access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
filesystem or document access, network or browser access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
filesystem or document access, network or browser access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
filesystem or document access, network or browser access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness