haoxuanjng-lang/kaggle-solutions-skills ↗★ 1
@haoxuanjng-lang/kaggle-solutions-skills
Evidence-backed Kaggle solution research: installable agent skill with offline knowledge
설치
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:haoxuanjng-lang/kaggle-solutions-skillskaggle-solutions-skills
Learn from past solutions. Turn evidence into a reusable agent skill.
Offline knowledge · Native DeepSeek Harness plugin · Multi-agent research · Ongoing knowledge review
Quick start · Workflow · DeepSeek plugin · Agent collaboration · Competition scenarios · Knowledge coverage · Research example · Ongoing updates
Kaggle solutions are scattered across discussions, notebooks, and author repositories. After finding a winning solution, you still need to assess its validation, decide what transfers to your data, combine research from several agents, and carry useful lessons into the next competition.
This project uses faridrashidi/kaggle-solutions as a discovery index. It distills author material that has actually been read into method cards with sources and conditions, then gives Codex and DeepSeek Harness access to the same knowledge base. The result: related solutions, validation risks, and comparable minimal experiments.
| 🔎 Find relevant solutions | 🧩 Research with multiple agents | 🌱 Build lasting knowledge |
|---|---|---|
| Search competitions by keyword and modality. Reuse the same knowledge through a Codex skill or native Harness tools. | Five roles—scout, evidence, validation, transfer, and synthesis—share sources and unknowns while preserving disagreements. | Review archive and method-card updates. Record actual experiments and counterexamples to improve reusable research knowledge. |
⚡ Quick start
1. Install with one command
Requires Node.js 20+. Install the complete skill from a public release without a GitHub token:
npx --yes --package=https://github.com/haoxuanjng-lang/kaggle-solutions-skills/releases/download/v0.2.1/haoxuanjng-lang-kaggle-solutions-skills-0.2.1.tgz kaggle-solutions-skills install
Add --force to update an existing installation, or --destination to choose a directory.
📦 GitHub Packages · Download the npm package · Download the skill ZIP
For full installation and release instructions, see Distribution.
Install from source / use the GitHub npm registry
Requires Python 3.10+. Offline search / show / patterns use only the standard library. The development installation also provides PyYAML for updates and validation.
git clone https://github.com/haoxuanjng-lang/kaggle-solutions-skills.git
cd kaggle-solutions-skills
python -m pip install -e .
python scripts/project.py install
The npm package is @haoxuanjng-lang/kaggle-solutions-skills, published at npm.pkg.github.com. After configuring registry authentication according to the official GitHub instructions:
npx --yes --registry=https://npm.pkg.github.com @haoxuanjng-lang/kaggle-solutions-skills@0.2.1 install
For ordinary downloads and installation, use the public release command above without configuring a registry. Python 3.10+ runs the skill's research tools; Node.js installs the files.
After installation, invoke the skill in a new chat:
$kaggle-solutions-skills Find similar problems and solutions for this competition, and propose experiments supported by evidence.
2. Search the knowledge base
The commands below run from the source checkout. For an installed skill, use the corresponding script paths under its installation directory.
# Inspect archive and research coverage
python skills/kaggle-solutions-skills/scripts/solutions.py stats
# Search for similar problems (medical imaging)
python skills/kaggle-solutions-skills/scripts/solutions.py search "medical imaging" --modality vision --limit 5
# Inspect solution leads for a competition
python skills/kaggle-solutions-skills/scripts/solutions.py show rogii-wellbore-geology-prediction --top-rank 5 --json
# Look up transferable methods (distillation)
python skills/kaggle-solutions-skills/scripts/solutions.py patterns "distillation" --json
3. Turn research into experiments
| Your goal | Example prompt |
|---|---|
| Compare historical approaches | $kaggle-solutions-skills Compare candidate retrieval and ranking in the OTTO winning solution, and propose a minimal controlled experiment. |
| Preserve competition lessons | $kaggle-solutions-skills Turn successes and failures from my experiment into method cards, preserving the actual results. |
| Update the knowledge base | $kaggle-solutions-skills Update the solution index, review upstream changes, and check the installed version. |
📖 Read the skill entry point · See the ROGII research handoff example
Installation directory and ZIP distribution
The default destination is $CODEX_HOME/skills/kaggle-solutions-skills, or ~/.codex/skills/kaggle-solutions-skills when the variable is unset. Use python scripts/project.py install --destination to choose a location.
Installation copies only the skill and offline knowledge. Caches, upstream checkouts, and private competition files stay in the source project.
python scripts/project.py package
The ZIP includes the skill, offline knowledge, MIT license, and upstream license. Packaging verifies CRC and file bytes. Published versions are also available from Releases.
🧭 Workflow
| Stage | Output | Evidence to preserve |
|---|---|---|
| Discover | Similar competitions and candidate solutions | Competition, links, rank leads, and upstream version |
| Read | Evidence from original material and code | Author statements, code locations, and access status |
| Distill | Method cards with conditions | Prerequisites, failure conditions, and source references |
| Hand off | Minimal controlled experiments | Target hypothesis, validation design, and resource limits |
| Feed back | New evidence and counterexamples | Actual results, runtime environment, and version |
For training, execution, and official Kaggle scores, pass the research to the user's chosen competition workflow. When agentic-kaggle-skill is already in use, follow it.
Evidence boundaries: Archive links, author reports, transfer hypotheses, local experiments, and official scores are recorded separately. Historical ranks help discover solutions; benefits in the target competition require actual validation.
🐋 Native DeepSeek Harness plugin
Current pain points and what the plugin helps you do
| Current pain point | How Harness helps | Deliverable |
|---|---|---|
| Many solutions are available, but their relevance to the current problem is unclear. | Search by problem, modality, and historical competition, then inspect solution leads. | Related competitions, author sources, and items to verify |
| Winning solutions contain many techniques; transfer can overlook validation and data boundaries. | Query method cards with conditions, then have evidence and validation agents investigate separately. | Preconditions, leakage risks, failure conditions, and minimal controls |
| Agents repeat searches, lose context, or mix author reports with actual experiments when merging results. | Generate separate task packets, a dependency graph, and source context for reviewer synthesis. | Research handoffs, unresolved disagreements, and experiment proposals |
Ask Harness to find similar problems for a new competition, compare solutions, organize a research team, or preserve feedback from actual experiments. Offline retrieval and task generation work directly. New model analysis uses your Harness configuration; experiments and official scores come from the competition execution workflow.
The same npm package includes the Cordis plugin, bundle patch, bilingual metadata, and tool icon. All six tools were successfully called in a local Harness 0.1.0-rc.6 full Web host, with the plugin manager showing the plugin mounted and enabled. The official 0.2.0-rc.2 tool runtime was also used to verify registration, execution, error propagation, and disposal. See the local validation record and screenshot.
Download the .tgz from the release, then run these commands in the directory containing it:
npx @deepseek-ai/dsh@0.2.0-rc.2 plugin --profile kaggle add ./haoxuanjng-lang-kaggle-solutions-skills-0.2.1.tgz
npx @deepseek-ai/dsh@0.2.0-rc.2 --profile kaggle --dump-config
| Retrieval tools | Evidence tools | Collaboration tools |
|---|---|---|
kaggle_solutions_search | kaggle_solutions_show | kaggle_research_scenarios |
kaggle_solutions_stats | kaggle_solutions_patterns | kaggle_research_plan |
Offline tools need no additional API key. Model conversations use Harness's own configuration. The plugin generates task packets; actual dispatch uses the host's subagent or workflow capabilities. Official Harness remains in developer preview. See the plugin guide for compatibility and workspace configuration.
Community discovery: listed in 1024Store (merged catalog PR). This is a community directory, not DeepSeek official certification. Use the verified Release tarball above for installation.
🤖 Multi-agent research
python skills/kaggle-solutions-skills/scripts/research_team.py plan --scenario otto --output workspaces/otto-research
python skills/kaggle-solutions-skills/scripts/research_team.py validate workspaces/otto-research
Each role receives its own task, pinned source context, result contract, and paths to prerequisite outputs. Evidence reading and validation review can run in parallel. Transfer analysis waits for both, and the reviewer combines traceable findings, disagreements, and minimal experiments.
In Codex, ask directly:
$kaggle-solutions-skills Use multiple agents to analyze OTTO retrieval and ranking, review validation leakage, and propose minimal transfer experiments. Keep missing target metrics unknown.
Generating task packets does not mean agents have run. After execution by the host, use validate --results to check role outputs and source references. A valid format still requires a researcher to assess whether sources support the claims. Collaboration workflow · Research usage guide
An actual OTTO research trial completed all five roles in dependency order. Independent evidence and validation agents ran in parallel, read cached author material, and contributed to a transfer experiment brief. This record validates the research workflow; competition gains remain unverified.
Archive facts now cite their own pinned source ID with archive_metadata; author observations keep their author source IDs. The validator rejects mixing these claim types. See the citation contract.
🏁 Research scenarios from real competitions
| Scenario ID | Original competition | Research focus | Minimal experiment direction |
|---|---|---|---|
otto | OTTO | Session boundaries, candidate retrieval, and ranking | Hold candidate generation fixed; compare only the ranking stage |
birdclef-2024 | BirdCLEF 2024 | Audio context, source quality, and pseudolabel isolation | Change one context or pseudolabel strategy with the same split and budget |
rogii | ROGII | Cross-well validation, alignment, and reliability routing | Hold the baseline fixed; compare one alignment or routing change |
amex | American Express | Customer isolation, historical features, and metric alignment | Preserve the customer split; add one group of historical statistical features |
m5 | M5 | Forecast horizons, recursive inference, and hierarchical metrics | Compare direct and recursive forecasts with the same horizon and training cost |
The five scenarios connect real competitions with author material recorded as read and with method cards. They support research and collaboration regression checks. These are historical solution research scenarios; this project has not reproduced their winning training pipelines or produced new official scores.
python skills/kaggle-solutions-skil