Skip to content
Skill

keyword-research-harvest

by zhongzhx

AI Summary

Use this skill when the user gives a topic or keyword set and wants a local literature-harvesting workflow, not a one-off manual search. This skill is self-contained. The target project does not need to already contain: The bundled copies live under:

Install

Copy this and paste it into Claude Code, Cursor, or any AI assistant:

I want to install the "keyword-research-harvest" skill in my project.

Please run this command in my terminal:
# Install skill into your project
mkdir -p .claude/skills/literature-harvest && curl --retry 3 --retry-delay 2 --retry-all-errors -o .claude/skills/literature-harvest/SKILL.md "https://raw.githubusercontent.com/zhongzhx/literature-harvest/main/SKILL.md"

Then restart Claude Code (or reload the window in Cursor) so the skill is picked up.

Description

Use when the user wants a reusable local workflow to search scholarly APIs for any topic keywords, build a candidate table, download accessible PDFs or HTML/XML full texts, run a second-pass HTML-to-PDF attempt, and deduplicate the downloaded files. Best for keyword-driven research-article harvesting that should be reproducible and saved into a project folder.

Keyword Research Harvest

Use this skill when the user gives a topic or keyword set and wants a local literature-harvesting workflow, not a one-off manual search.

What this skill does

• Bundles its own literature_harvest/scripts stack, including search_pubmed.py, search_europepmc.py, search_crossref.py, search_openalex.py, merge_and_deduplicate.py, download_fulltexts.py, and harvest_utils.py. • Searches PubMed/PMC, Europe PMC, Crossref, and OpenAlex through the bundled pipeline. • Builds a no-dedup candidate table first. • Downloads all legally accessible files to a new run folder. • Saves HTML/XML when PDF is not directly available. • Runs a second pass to chase PDF links from saved HTML pages. • Produces a deduplicated file folder and manifest after downloading.

Dependency model

This skill is self-contained. The target project does not need to already contain: • literature_harvest/ • search_pubmed.py • download_fulltexts.py The bundled copies live under: • literature_harvest/scripts/

Files in this skill

• scripts/run_keyword_harvest_no_dedup.py Use this to launch a new broad keyword harvest run into a new run folder. • scripts/continue_download_and_dedup.py Use this to resume pending downloads, try HTML-to-PDF second pass, and build a deduplicated download set. • literature_harvest/scripts/ Bundled search and download dependencies copied from a working local harvest stack so the skill is portable. • references/config_template.json Copy and edit this for the topic-specific query set and filtering terms. • references/prompt_template.md Reusable prompt for another AI/agent.

Discussion

0/2000
Loading comments...

Health Signals

MaintenanceCommitted 3mo ago
Stale
AdoptionUnder 100 stars
68 ★ · Niche
DocsREADME + description
Well-documented

GitHub Signals

Stars68
Forks3
Issues0
Updated3mo ago
View on GitHub
MIT License

My Fox Den

Community Rating

Sign in to rate this booster

Works With

Claude Code