Command-line GitHub repository miner that aggregates dataset for your researches
-
Updated
Sep 17, 2026 - JavaScript
Command-line GitHub repository miner that aggregates dataset for your researches
GitHub repositories dataset that contains sample repositories (SRs), with their metrics and metadata
Large-scale dataset of 65,195 Amazon Alexa skills: metadata, descriptions, permissions, and privacy policy links for privacy and software-engineering research (ADMA 2021)
Curated 2026 conference papers on Claude Code and Codex CLI—benchmarks, methods, failures, exact models, and auditable evidence.
By providing standard data sets for orchestration and scheduling of computing resources, it can help researchers better study and train scheduling algorithms.
SARMF – Smart Contract Automated Remediation and Mitigation Framework (DOI-backed reproducible smart contract security engineering methodology)
A curated, citation-backed, confidence-tier-annotated dataset of human gestational physiology, distributed as a Python package (pip install nidus) with an interactive Streamlit dashboard. JSON + JSON Schema, MIT/CC-BY-4.0, Zenodo-deposited.
Standardized parameter ontology and open dataset for Microbial Electrochemical Systems (MES) research — 704 parameters across 13 categories, 23,568 papers indexed. Public mirror.
Pipeline for building a research dataset of world-model projects from GitHub, including relevance filtering, technical characterisation, and application-domain analysis.
Public Eisenhardt case-study reference map and reproducible Codex workspace-maintenance toolkit
Read-only MCP server for anonym.community's privacy research corpus — 1,478 PII pain points, 98 structural drivers, 240 jurisdictions, 134 FAQs, 1,600+ papers.
Racial equity plans, accountability scores, and disparity metrics for the 20 largest U.S. cities (April 2026)
A curated evidence atlas of 15 Building Simulation papers on AI-based energy prediction, control, design and uncertainty.
Data package for the questionnaire from the DP-Next project
Tracing historical stablecoin blacklisting, dataset and dashboard.
Research datasets and experimental results from comprehensive ML benchmarking on AMD MI300X hardware. Contains raw telemetry, processed metrics, and reproducible analysis artifacts from both inference and training workloads, supporting research and comparative studies.
Public landing page and documentation for the SPIDER synthetic person information dataset for entity resolution.
Curated, versioned index of 403 repositories for acoustics, audio, musical instruments, organology, spatial audio, and synthesis.
Tracing historical stablecoin blacklisting, dataset and dashboard.
Living evidence map for ROS 2, VLA/VLM/LLM, Edge AI, and Sim2Real research.
To associate your repository with the research-dataset topic, visit your repo's landing page and select "manage topics."