SC

scholar-evaluation

Evaluates scholarly research using the ScholarEval framework with quantitative scoring.

Install

mkdir -p .claude/skills/scholar-evaluation && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/1533" && unzip -o skill.zip -d .claude/skills/scholar-evaluation && rm skill.zip

Installs to .claude/skills/scholar-evaluation

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing with quantitative scoring and actionable feedback.
255 charsno explicit “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Intermediate

Key capabilities

  • Systematically evaluate scholarly work using the ScholarEval framework
  • Assess research quality across dimensions like methodology and analysis
  • Provide quantitative scoring on a 5-point scale
  • Generate actionable feedback for academic improvement
  • Integrate with scientific-schematics for visual communication

How it works

The skill applies the ScholarEval framework to assess research quality dimensions, utilizing a retrieval-augmented approach to ground evaluations in peer-reviewed criteria.

Inputs & outputs

You give it
Scholarly work such as a research paper, proposal, or literature review
You get back
Structured evaluation report with dimension-based scores and actionable recommendations

When to use scholar-evaluation

  • Evaluate a research paper's methodology
  • Score scholarly work quality
  • Assess research writing effectiveness

About this skill

Scholar Evaluation

Purpose

Provide developmental, evidence-traceable feedback on a scholarly work: paper, draft, protocol, literature synthesis, or research idea. Use qualitative judgment first. Optional scores only describe how submitted evidence maps to a predeclared bounded rubric.

This skill also audits whether a low-stakes assessment process documents its construct, provenance, rater quality, uncertainty, traceability, sensitivity, fairness, accessibility, privacy, and human governance.

Hard safety boundary

Never use this skill to automate, recommend, materially influence, or score:

  • hiring, promotion, or tenure;
  • admissions;
  • grants or other funding;
  • prizes, honors, or awards;
  • discipline, dismissal, or sanctions; or
  • any other high-impact personnel decision.

Never rank people. Never reduce a person to a composite score. Never infer ability, character, integrity, protected traits, future performance, or worth. A nominal human-in-the-loop does not remove this boundary.

If asked for a prohibited use, stop. Offer developmental comments on a scholarly work or a process-only audit that does not process applications, compare people, recommend an outcome, or advise a decision.

Do not issue publication-readiness, accept/reject, or “top-tier” judgments.

Read references/responsible_assessment.md before any organizational use.

ScholarEval status

The referenced ScholarEval project is an experimental literature-grounded research-idea evaluation framework, not validated psychometrics.

The verified primary record is Moussa et al., ScholarEval: Research Idea Evaluation Grounded in Literature, arXiv:2510.16234v2, revised 2026-02-28. It reports a retrieval-augmented soundness/contribution framework, a 117-idea four-discipline dataset, coverage experiments, and a user study.

Do not generalize those results to person assessment, consequential decisions, all disciplines, or this skill's rubric. No peer-reviewed publication status was verified during the dated review. See references/source_ledger.md.

Metric and prestige policy

Do not score or infer quality from:

  • Journal Impact Factor or other journal measures;
  • h-index, publication counts, or citation counts;
  • altmetrics or attention;
  • journal, conference, venue, institution, employer, or geographic prestige;
  • author affiliation, reputation, network, or career path.

The rubric validator rejects common proxy-measure criteria.

If a qualified reviewer mentions an indicator descriptively outside the scoring tools, record its exact purpose, source, coverage, field and time effects, uncertainty, missingness, biases, gaming risk, and why it does not directly measure quality. Never hide indicators inside an opaque composite.

Data boundary

Bundled scripts accept only strict local JSON/CSV containing pseudonymous IDs, bounded ratings, statuses, uncertainty, and local references.

Do not put raw private applications, CVs, letters, reviewer identities, contact details, protected attributes, or source-document text in inputs, outputs, logs, examples, or prompts. Keep source content in the authorized records system and use opaque local references.

Allowed classifications are:

  • synthetic
  • public_scholarly_work
  • deidentified_low_stakes

No script searches the web, loads environment files, reads credentials, calls a model, executes supplied text, deserializes executable objects, or launches a process.

Use Bash only to invoke the documented local python3 commands.

Workflow

1. Confirm allowed use and authorization

Record:

  • developmental purpose;
  • unit of assessment: scholarly_work;
  • work type, stage, discipline, language, and audience;
  • authorized source location and data classification;
  • accountable committee owner;
  • conflicts and recusals;
  • accessibility and accommodation process;
  • appeal or correction route; and
  • data purpose, access, retention, and deletion.

Stop on a prohibited decision context or unnecessary private data.

2. Define the construct before criteria

State:

  • what quality or support is being examined;
  • excluded constructs;
  • intended interpretation;
  • contexts where the interpretation does not travel;
  • evidence requirements; and
  • known limitations.

Start with values and disciplinary context, not available metrics.

3. Adapt and validate the rubric

Begin with assets/rubric_template.json, then obtain qualified disciplinary, assessment-methods, stakeholder, accessibility, privacy, and fairness review.

The template deliberately records content validity as not_established. Do not change that status without documented evidence for the exact intended use.

Validate structure:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/validate_rubric.py \
  --rubric assets/rubric_template.json

Read references/evaluation_framework.md for construct, anchor, validity, and rater guidance.

4. Build traceable evidence records

Reviewers may read an authorized work outside the scripts. Record only stable local locators and claim references in assets/evidence_manifest_template.json.

For every criterion, distinguish:

  • observed evidence from interpretation;
  • supporting from contrary evidence;
  • available from unavailable evidence;
  • missing from not_applicable; and
  • uncertainty from absence.

Failure to find prior work does not prove novelty.

5. Rate independently

Use assets/evaluation_template.json. Each criterion must be:

  • rated with an anchor score, bounded uncertainty, evidence IDs, and a local rationale reference;
  • missing with null score/uncertainty and a rationale reference; or
  • not_applicable with null score/uncertainty and a rationale reference.

Do not encode missing or not-applicable as zero. Raters should train, calibrate, disclose conflicts, rate independently, and document disagreement.

6. Run local quality checks

Bounded scoring, without labels or recommendation:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/calculate_scores.py \
  --rubric assets/rubric_template.json \
  --evaluation assets/evaluation_template.json

Evidence traceability:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_traceability.py \
  --rubric assets/rubric_template.json \
  --evaluation assets/evaluation_template.json \
  --evidence assets/evidence_manifest_template.json

Inter-rater agreement:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/summarize_agreement.py \
  --rubric assets/rubric_template.json \
  --ratings assets/ratings_template.csv

Weight sensitivity requires two or more distinct scholarly-work evaluation files:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/weight_sensitivity.py \
  --rubric assets/rubric_template.json \
  --evaluation /tmp/work-a-evaluation.json \
  --evaluation /tmp/work-b-evaluation.json

Process controls:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_process.py \
  --process assets/process_checklist_template.json

The checklist template is intentionally unconfirmed and fails closed. Instructions and exact schemas are in references/local_tooling.md.

7. Synthesize qualitative findings

Lead with criterion-level evidence, not the composite. For each criterion:

  1. cite evidence references;
  2. state rated, missing, or not_applicable;
  3. explain the anchor interpretation;
  4. report score and uncertainty only if rated;
  5. note disagreements and context;
  6. identify strengths and limitations; and
  7. offer non-prescriptive improvement options.

Generate an empty-reference scaffold if useful:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/generate_report_scaffold.py \
  --rubric assets/rubric_template.json \
  --evaluation assets/evaluation_template.json \
  --output /tmp/developmental-report-scaffold.json

The scaffold does not read source documents or draft findings.

8. Human review and release

Before releasing an organizational report, a qualified accountable human committee must verify:

  • construct and rubric provenance;
  • content-validity evidence and limits;
  • rater training, agreement, inter-rater reliability evidence, and drift;
  • evidence traceability and source access;
  • missingness, not-applicable rationales, and uncertainty;
  • weight sensitivity and order instability;
  • disciplinary and subgroup bias review;
  • conflicts and recusals;
  • accessibility and accommodations;
  • privacy, minimization, retention, and output controls; and
  • correction or appeal information.

Document dissent. Do not imply consensus, validity, or precision beyond the evidence. Periodically evaluate the evaluation and retire harmful criteria.

Interpretation rules

  • A score is an ordinal rubric summary, not a natural measurement.
  • Normalization does not repair incomplete evidence.
  • The bundled uncertainty range is not a confidence interval.
  • Agreement does not establish reliability, validity, fairness, or correctness.
  • Stable results under tested weights do not establish validity.
  • The overall score never overrides criterion evidence or qualified judgment.
  • No output is a decision recommendation.

Bundled resources

  • references/responsible_assessment.md — safety, metrics, governance, accessibility, privacy, and bias.
  • references/evaluation_framework.md — ScholarEval boundary, construct, criteria, anchors, validity, and interpretation.
  • references/local_tooling.md — strict schemas, formulas, commands, and output behavior.
  • references/source_ledger.md — authoritative sources and publication-status verification dated 2026-07-23.
  • references/security_validation.md — baseline remediation, validation, and residual security-scan record.
  • assets/rubric_template.json — bounded rubric template.
  • assets/evaluation_template.json — rating template.
  • assets/evidence_manifest_template.json — traceability template.
  • assets/process_checklist_template.json — fail-closed process checklist.
  • assets/ratings_template.csv — syn

Content truncated.

When not to use it

  • When the work is outside the scope of scholarly or research writing
  • When domain-specific expertise is required that the framework lacks

Prerequisites

OPENROUTER_API_KEY

Limitations

  • Some dimensions may not apply to all work types, such as data collection for theoretical papers
  • The framework complements but does not replace domain-specific expertise

How it compares

Unlike manual peer review, this skill provides a standardized, quantitative, and dimension-specific evaluation framework that can be applied consistently across different research types.

Compared to similar skills

scholar-evaluation side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
scholar-evaluation (this skill)42moReviewIntermediate
scientific-critical-thinking187moReviewAdvanced
tooluniverse-drug-research32moNo flagsAdvanced
literature-review5592moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by K-Dense-AI

View all by K-Dense-AI

literature-review

K-Dense-AI

Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).

5591,298

markitdown

K-Dense-AI

Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. Use when converting documents to markdown, extracting text from PDFs/Office files, transcribing audio, performing OCR on images, extracting YouTube transcripts, or processing batches of files. Supports 20+ formats including DOCX, XLSX, PPTX, PDF, HTML, EPUB, CSV, JSON, images with OCR, and audio with transcription.

177310

scientific-writing

K-Dense-AI

Write scientific manuscripts. IMRAD structure, citations (APA/AMA/Vancouver), figures/tables, reporting guidelines (CONSORT/STROBE/PRISMA), abstracts, for research papers and journal submissions.

94309

exploratory-data-analysis

K-Dense-AI

Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats. This skill should be used when analyzing any scientific data file to understand its structure, content, quality, and characteristics. Automatically detects file type and generates detailed markdown reports with format-specific analysis, quality metrics, and downstream analysis recommendations. Covers chemistry, bioinformatics, microscopy, spectroscopy, proteomics, metabolomics, and general scientific data formats.

15114

infographics

K-Dense-AI

Create professional infographics using Nano Banana Pro AI with smart iterative refinement. Uses Gemini 3 Pro for quality review. Integrates research-lookup and web search for accurate data. Supports 10 infographic types, 8 industry styles, and colorblind-safe palettes.

1141

pptx-posters

K-Dense-AI

Create research posters using HTML/CSS that can be exported to PDF or PPTX. Use this skill ONLY when the user explicitly requests PowerPoint/PPTX poster format. For standard research posters, use latex-posters instead. This skill provides modern web-based poster design with responsive layouts and easy visual integration.

911

You might also like

scientific-critical-thinking

davila7

Evaluate research rigor. Assess methodology, experimental design, statistical validity, biases, confounding, evidence quality (GRADE, Cochrane ROB), for critical analysis of scientific claims.

1888

tooluniverse-drug-research

mims-harvard

Generates comprehensive drug research reports with compound disambiguation, evidence grading, and mandatory completeness sections. Covers identity, chemistry, pharmacology, targets, clinical trials, safety, pharmacogenomics, and ADMET properties. Use when users ask about drugs, medications, therapeutics, or need drug profiling, safety assessment, or clinical development research.

323

literature-review

K-Dense-AI

Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).

5591,298

openalex-database

davila7

Query and analyze scholarly literature using the OpenAlex database. This skill should be used when searching for academic papers, analyzing research trends, finding works by authors or institutions, tracking citations, discovering open access publications, or conducting bibliometric analysis across 240M+ scholarly works. Use for literature searches, research output analysis, citation analysis, and academic database queries.

48202

market-research-reports

davila7

Generate comprehensive market research reports (50+ pages) in the style of top consulting firms (McKinsey, BCG, Gartner). Features professional LaTeX formatting, extensive visual generation with scientific-schematics and generate-image, deep integration with research-lookup for data gathering, and multi-framework strategic analysis including Porter's Five Forces, PESTLE, SWOT, TAM/SAM/SOM, and BCG Matrix.

38162

annas-archive-ebooks

ratacat

Use when needing to look up book content, find a book by title/author, download an ebook, or reference material from a published book. Triggers on book lookups, ebook downloads, "find the book", "get the PDF/EPUB of". Downloads produce PDF/EPUB/MOBI files - use ebook-extractor skill to convert to text.

22177

Search skills

Search the agent skills registry