Heureka Skills

Skills for scientific discovery

A growing library of agent skills for real research work — pulling data, running models, analyzing and communicating results, and drafting what funders ask for. Every skill is reviewed by scientists before it’s published.

Install into Heureka Bench in one click, or ask your agent to find and install one mid-conversation. Every skill is served from a public repo, readable by anything.

gaudi analysis

Unsupervised multi-omics integration with GAUDI — two-stage UMAP, HDBSCAN clustering, and SHAP metagenes, in R or Python. Use when two or more omics layers share samples and you need sample clusters plus the features driving them.

Heureka Labs · v1.3.0 🥇 3 installs 4 files
alphafold data

Retrieve predicted protein structures from the AlphaFold database by UniProt accession, and read their confidence metrics correctly — per-residue pLDDT, the PAE matrix for multi-domain proteins, and AlphaMissense variant annotations.

Heureka Labs · v1.3.0 🥈 2 installs 1 file
biopython utility

Molecular biology toolkit with Biopython — sequence manipulation, FASTA/GenBank/PDB parsing, phylogenetics, BLAST automation, and programmatic NCBI/PubMed access via Bio.Entrez.

K-Dense Inc. (adapted by Heureka Labs) · v1.4.0 🥉 1 install 8 files
boltz2-nim models

Predict biomolecular structure and binding affinity with Boltz2 NIM — protein, protein-ligand, DNA and RNA complexes from SMILES or CCD ligands, returning mmCIF and pIC50 scores, via the hosted NVIDIA API or local Docker.

NVIDIA BioNeMo Agent Toolkit (adapted by Heureka Labs) · v1.4.0 1 install 6 files
esm models

Run the ESM protein models — ESMFold2 all-atom structure and complex prediction with DNA, RNA, ligands and modified residues, ESMC embeddings up to 6B, ESM3 generation and inverse folding. Covers which weights are open and under what licence, the CUDA, Apple Silicon and hosted routes, and the dependency pin that breaks upstream's own quickstart.

K-Dense Inc. (adapted by Heureka Labs) · v1.4.0 1 install 6 files
scikit-bio analysis

Biological data analysis with scikit-bio — sequence handling, alignments, phylogenetic trees, alpha and beta diversity including UniFrac, ordination (PCoA), and PERMANOVA. Built for microbiome work.

K-Dense Inc. (adapted by Heureka Labs) · v1.2.0 1 install 2 files
anndata utility

Work with AnnData annotated matrices and .h5ad files — layers, obs/var metadata, sparse backing, on-disk access, and format conversion. The data-format layer under the single-cell ecosystem.

K-Dense Inc. (adapted by Heureka Labs) · v1.3.0 1 install 6 files
autodock-vina models

Molecular docking with AutoDock Vina — receptor and ligand preparation, grid box definition, pose prediction and scoring. Use to predict how a small molecule binds a protein, to redock a known complex for validation, or to screen a compound set.

Heureka Labs · v1.3.0 1 install 3 files
cellxgene-census data

Query the CZ CELLxGENE Census for versioned public single-cell and spatial transcriptomics data — cell metadata, expression slices, summary counts, source H5AD downloads, and embeddings across organisms, tissues, diseases, and cell types.

K-Dense Inc. (adapted by Heureka Labs) · v1.2.0 1 install 4 files
metabolights data

Find and retrieve public metabolomics studies from EMBL-EBI MetaboLights — search by disease, tissue or organism, read the ISA-Tab study design and factors, and download the Metabolite Assignment File holding identified metabolites and their per-sample intensities. Covers LC-MS, GC-MS and NMR studies across human, mouse and rat.

Heureka Labs · v1.0.0 1 install 1 file
pathway-cca-coessentiality analysis

Score pathway-pathway association by PCA to k components then first canonical correlation

Pol Castellano-Escuder · v1.3.0 1 install 1 file
pathway-enrichment analysis

Pathway and gene-set enrichment on a gene list or ranked table — over-representation (Enrichr, g:Profiler), preranked GSEA, and ssGSEA/GSVA against GO, KEGG, Reactome, and MSigDB, with background choice and FDR handled correctly.

K-Dense Inc. (adapted by Heureka Labs) · v1.3.0 1 install 4 files
pydeseq2 analysis

Differential gene expression for bulk RNA-seq with PyDESeq2 — formulaic designs, Wald tests, FDR correction, LFC shrinkage, and result visualization.

K-Dense Inc. (adapted by Heureka Labs) · v1.1.0 1 install 5 files
scvi-tools models

Deep generative models for single-cell omics with scvi-tools — probabilistic batch correction (scVI), transfer learning, differential expression with uncertainty, and multi-modal integration (totalVI, MultiVI).

K-Dense Inc. (adapted by Heureka Labs) · v1.4.0 1 install 9 files
statistical-power utility

Sample-size and power calculations for planning a study — a priori power, minimum detectable effect, and power curves. Closed-form for t-tests, ANOVA, proportions, correlation, and regression; Monte Carlo simulation for the rest.

K-Dense Inc. (adapted by Heureka Labs) · v1.2.1 1 install 4 files
4dn data

Search and download 4D Nucleome chromatin architecture data — in situ Hi-C, Micro-C, HiChIP, ChIA-PET, SPRITE, DamID, Repli-seq, ATAC-seq and imaging (DNA FISH, single particle tracking) across human and mouse cell lines and tissue. Walks ExperimentSet to Experiment to File, checks each file's access tier before transferring, and pulls open files from the AWS Open Data bucket with no account.

Heureka Labs · v1.0.0 1 file
all-of-us data

Browse the All of Us public Data Browser API anonymously — OMOP domains, concept and participant counts, survey modules and questions, and aggregate genomic allele frequencies — and report what Researcher Workbench access requires. Delivers a catalogue of what is measured, never participant data, which never leaves the Workbench.

Heureka Labs · v1.0.1 1 file
bindcraft models

Design de novo protein binders against a chosen epitope with BindCraft — trimming the target, hotspot syntax, the settings and filter file grammar, what its accept/reject thresholds actually compare, and which interface metrics the PyRosetta-free FreeBindCraft build replaces with constants. The design stage needs an NVIDIA GPU.

Heureka Labs · v1.0.0 1 file
binder-design-filtering analysis

Rank and cut de novo protein binder designs before paying to synthesise them — interface confidence from the PAE matrix (ipSAE, ipTM, pDockQ), pLDDT scoped to the binder chain, self-consistency by DockQ, interface geometry and sequence liabilities. Every threshold carries its source, and the filters are scored against a public labelled benchmark of 402 designs.

Heureka Labs · v1.0.0 1 file
bioagent-bench data

Access BioAgent Bench, ten end-to-end bioinformatics pipeline tasks where an agent is handed raw inputs and a written objective and must return one deliverable file. Fetch the task index and per-task inputs from an open OSF deposit that needs no account, with the truth files projected out on load. Covers RNA-seq, single-cell, variant calling, metagenomics and comparative genomics.

Heureka Labs · v1.0.0 1 file
biomedarena data

BioMedArena is an MIT harness that puts 166 registered benchmark entries — 155 biomedical benchmarks plus 11 aliases — behind one CLI. This skill maps that index, resolves every entry to the dataset the loader actually reads, and measures the licence and access gate on each, because the harness licence covers none of them. The map lands on disk as CSV.

Heureka Labs · v1.0.0 1 file
biomnibench-da data

Access BiomniBench-DA, 50 publicly released biomedical data-analysis tasks built from Nature, Cell and Science papers and graded on the whole agent trajectory against expert rubrics. Browse the task catalogue and trace a task back to its source paper without an account, then pull one task's real data with a Hugging Face login. Leaves the rubrics alone.

Heureka Labs · v1.0.0 1 file
biostudies data

Search and retrieve studies from EBI BioStudies, the archive holding the ArrayExpress collection of functional-genomics submissions whose E-MTAB accessions mostly never reach GEO. Resolve an accession, traverse the nested section tree that carries title, organism and protocols, and download processed files and MAGE-TAB sample tables.

Heureka Labs · v1.1.0 1 file
bixbench data

Access BixBench, a dataset of 205 open-ended bioinformatics questions paired with 59 data capsules taken from published papers' analysis notebooks. Load the index, browse questions by category and source paper, and pull a capsule's real data onto disk. A benchmark by origin, and a browsable corpus of worked analysis cases in its own right.

Heureka Labs · v1.2.0 1 file
bixbench3 data

Access BixBench3, a benchmark of 20 research-study-scale computational biology tasks in which an agent must rebuild a published paper's analysis from its raw data. Fetch the release index, per-task manifests carrying per-file provenance and checksums, and study data from open buckets that need no account. The task prompts themselves sit behind a free click-through.

Heureka Labs · v1.0.0 1 file
boltzgen models

Design protein and peptide binders against a target with BoltzGen — the design specification YAML, the six protocols, weight provenance, and how to read the design confidence metrics. Generation, not prediction. Needs an NVIDIA GPU to run the model; the specification, install and output checks run on CPU.

Heureka Labs · v1.0.0 1 file
brain-researcher analysis

Audit your own neuroimaging analysis with Brain Researcher — seal a commitment before you look, grade the evidence, bound the claim by how it survived pipeline variation, and re-hash the record afterwards. Covers the 10 versioned MCP contracts, the 1,693-dataset access catalogue, and the pinned open-data path. The BR-KG graph itself is private and is not shipped.

Brain Researcher Team (adapted by Heureka Labs) · v1.0.0 1 file
caliby models

Design protein sequences against a structural ensemble rather than one backbone with Caliby — a Potts model whose parameters average across conformers, so one sequence is optimised for all of them at once. Covers where ensembles come from, what the energy means, and the silent ways the molecule Caliby designs stops being the molecule you handed it.

Heureka Labs · v1.0.0 1 file
cellprofiler analysis

Run CellProfiler pipelines headlessly and edit them as text — the -c -r invocation, LoadData CSVs that carry per-image metadata, batch groups, illumination correction before measurement, and joining the exported per-object tables back to the images they came from.

Heureka Labs · v1.0.0 1 file
codex-phenocycler analysis

Take a CODEX or PhenoCycler image from pixels to a cell-by-marker table — choosing segmentation channels, running Cellpose over a multiplex stack, measuring per-cell intensity in micrometres, and building the AnnData that spatial analysis starts from. Includes the licence and token gates on the tools this workflow is usually built with.

Heureka Labs · v1.0.0 1 file
cross-modal-registration analysis

Bring a fluorescence, CODEX or IHC section into the same coordinate frame as its H&E image using ACCREDIT — building one representation both modalities share, searching orientation and scale before refining, and reading the reference-free quality score honestly enough to separate a good registration from a wrong one that scores well.

Heureka Labs · v1.1.0 1 file
datamol analysis

Cheminformatics with datamol, a Pythonic layer over RDKit with sensible defaults — SMILES parsing and standardization, descriptors, fingerprints, clustering, 3D conformers, and parallel processing.

K-Dense Inc. (adapted by Heureka Labs) · v1.3.0 9 files
dbgap data

Find controlled-access human cohorts in the dbGaP study catalogue by disease, cohort name or assay, and build a table on disk of accessions, versions, study design, subject counts, consent codes and their data use limitations, and the access policy. Delivers study metadata only — the individual-level genotypes and phenotypes stay behind an application this skill cannot make.

Heureka Labs · v1.1.0 1 file
depmap data

Retrieve DepMap cancer dependency data pinned to a named quarterly release — CRISPR gene effect (Chronos), expression, copy number, mutations, drug sensitivity and cell-line annotation across ~1,178 models. Resolves a release name to a citable, versioned figshare record, because gene effect scores move between quarters and an unpinned analysis is not reproducible.

Heureka Labs · v1.0.0 1 file
diffdock-nim models

Run DiffDock molecular docking via NVIDIA NIM to predict small-molecule binding poses against protein targets — blind ligand docking of SMILES or SDF inputs, returning ranked poses with confidence scores, hosted NVIDIA API or local Docker.

NVIDIA BioNeMo Agent Toolkit (adapted by Heureka Labs) · v1.4.0 6 files
dryad data

Retrieve published datasets from Dryad by DOI — the CC0 data behind a paper with the depositor's methods prose, funder and ROR affiliations, spatial coverage and linked publication. Resolve a paper DOI to its deposit and page a file manifest with SHA-256 digests anonymously, then write a dataset card. File bytes need a free Dryad API token.

Heureka Labs · v1.0.0 1 file
ena data

Find and download public sequencing runs from the European Nucleotide Archive — search 44 million runs by study, organism, instrument, library strategy or date, then get every FASTQ URL together with its exact size and MD5 in the same call, so the transfer can be totalled and refused or subset before a byte moves. Resolves SRA and DDBJ accessions too.

Heureka Labs · v1.0.0 1 file
experimental-design utility

Design a study before data is collected — pick a design, randomize, block and stratify, and lay out treatment combinations. Covers factorial and fractional-factorial DOE, crossover, split-plot, Latin squares, and plate layouts.

K-Dense Inc. (adapted by Heureka Labs) · v1.3.0 5 files
genie3 models

Generate de novo protein backbones with Genie 3, the SE(3)-equivariant diffusion model from the AlQuraishi lab — unconditional folds, motif scaffolding and target-conditioned binders, on a laptop CPU with no GPU. Covers what the written PDB really contains (a Cα trace, not the all-atom structure the name implies) and the geometry checks that catch a silently exploded sample.

Heureka Labs · v1.0.0 1 file
geo data

Find, triage and download public gene expression datasets from NCBI GEO — search series by tissue, disease, organism, assay and file type through E-utilities, read the esummary triage fields before transferring anything, then pull the counts matrices from all three places GEO keeps them — suppl/, the per-platform series matrix, and the per-sample directories.

Heureka Labs · v1.3.0 1 file
graphical-abstract communication

Turn a paper or results summary into an accurate, accessible graphical abstract as editable SVG — choose the figure type, compose it, then audit palette contrast, greyscale separation and colour-vision safety before submission.

Heureka Labs · v1.1.0 4 files
gtex data

Query GTEx human tissue expression and eQTLs through the GTEx Portal API — median TPM across 54 tissues including heart left ventricle, atrial appendage and skeletal muscle, per-sample values by donor age bracket, and significant cis-eQTLs by gene and tissue. Resolves gene symbols to the versioned GENCODE ids GTEx requires, and says which data is open and which is dbGaP-controlled.

Heureka Labs · v1.1.0 1 file
gwas-catalog data

Retrieve variant-trait associations from the NHGRI-EBI GWAS Catalog REST API — by trait, study accession, variant or gene — together with the discovery and replication cohort ancestry, sample sizes and effect estimates that decide whether an association transfers to anyone else. Handles the HAL response shape, the 20-row default page, and p-values that underflow a float to 0.0.

Heureka Labs · v1.0.0 1 file
hubmap data

Query HuBMAP — the NIH atlas of healthy human tissue — for datasets, donors, samples and files across kidney, lung, placenta, spleen, heart, intestine and twenty more organs. Covers CODEX, MIBI, Visium, snRNA-seq and mass-spec assays, the donor to sample to dataset provenance chain, and which assays are openly downloadable versus protected human sequence.

Heureka Labs · v1.0.0 1 file
idr data

Browse and fetch public imaging studies from the Image Data Resource — resolve a published idr0000 accession to its OMERO screens and projects, walk plates, wells and datasets down to single images, read per-image gene, phenotype and pixel-size annotations, check each study's own licence, and pull OME-Zarr arrays or the EBI mirror.

Heureka Labs · v1.0.0 1 file
impc data

Query IMPC / KOMP2 mouse knockout phenotypes over the EBI Solr API — Mammalian Phenotype calls for ~8,000 knocked-out genes with p-values, effect sizes, zygosity, sex and phenotyping centre, plus the separate viability and fertility screens that say whether the homozygote survives. Resolves a human gene to its mouse ortholog, and distinguishes never tested from tested and normal.

Heureka Labs · v1.0.0 1 file
journal-editor-finder utility

Find which of a journal's handling editors have recently handled papers closest to a manuscript, using only the published record — ranks named editors by the papers they actually took through review at that journal, excludes co-authors, flags institutional overlap, and reports plainly when a journal names no handling editor anywhere public.

Heureka Labs · v1.0.0 1 file
journal-selection utility

Build a shortlist of candidate journals for a finished manuscript, grouped by the angle the author chooses to lead with — the subject framing and the methods framing are found separately, and each angle gets its own specialist and broad-scope options with the evidence behind them.

Heureka Labs · v1.0.0 1 file
lab-bench data

Fetch and cite LAB-Bench and LAB-Bench 2, the benchmarks for AI agents doing biology research — literature, figures, tables, databases, protocols, sequences, cloning, patents, clinical trials. Per-subset counts, the record schema, the standard citation, the share-alike constraint that governs reuse, and how to avoid contaminating the benchmark.

Heureka Labs · v1.0.0 1 file
latch-bench data

Access the LatchBio agent benchmarks — scBench, SpatialBench, EpiBench, VariantBench and two long-horizon variants. Fetch the 45 publicly released evaluations, read the deterministic grader contract declared inside each task file, and pull the AnnData analysis snapshots a task starts from. The full benchmarks are withheld upstream; this reaches the public sample.

Heureka Labs · v1.0.0 1 file
mdarena data

Access MDArena, 50 containerised molecular dynamics tasks — system preparation, parameterization, trajectory analysis, free energy, enhanced sampling — used to score coding agents. Map the catalogue and the per-task resource contract, check the force-field and PDB provenance, and find which 43 of the 50 run from a bare clone. Reads metadata only and never task content.

Heureka Labs · v1.1.0 1 file
metabolomics-workbench data

Find and retrieve public metabolomics and lipidomics studies from the NIH Metabolomics Workbench — study design, experimental factors, named metabolites and measured values over its path-segment REST API — plus the RefMet endpoints that map arbitrary metabolite names onto one nomenclature so two studies can be joined.

Heureka Labs · v1.0.0 1 file
microscopy-quantification analysis

Quantify fluorescence microscopy images — open acquisition formats with pixel size and channels intact, segment nuclei or cells, and export per-object area, shape, intensity and colocalization as a tidy CSV with QC overlays.

Heureka Labs · v1.1.0 1 file
motrpac data

Retrieve MoTrPAC multi-omics from the openly licensed R data package and public bucket instead of the account-gated Data Hub API — transcriptomics, proteomics, phosphoproteomics, acetylome, metabolomics, ATAC-seq and RRBS across twenty rat tissues including heart, skeletal muscle, liver and adipose, in endurance-trained animals. Young-adult 6-month cohort, not an aged one.

Heureka Labs · v1.1.0 1 file
multiplex-imaging-io utility

Open multiplex imaging files — Akoya QPTIFF from PhenoCycler and Vectra, CODEX multichannel TIFF, OME-TIFF, and folders of single-channel TIFFs — resolving channels by marker name rather than position, reading regions instead of whole slides, and recovering pixel size in micrometres from the file's own tags.

Heureka Labs · v1.0.0 1 file
nih-dms-plan grants

Draft an NIH Data Management and Sharing plan from a project's aims, in the 2026 structured format required for applications since May 2026 and for all awards from FY2027. Produces a clean plan plus a separate summary of what was proposed, inferred, and left for the PI. Covers repository choice, the word limits, the genomic data sharing elements, and the Just-in-Time and RPPR transition paths.

Heureka Labs · v1.0.0 2 files
nsf-biosketch grants

Stage NSF's Biographical Sketch Common Form for one senior person in SciENcv's own entry order and audit the open record behind their ORCID iD — products deduplicated across Crossref and DataCite so software and datasets are counted, ranked against the proposed project, plus a summary naming every field only the researcher can supply. SciENcv generates the PDF; this does not.

Heureka Labs · v1.0.0 1 file
nsf-coa grants

Fill NSF's Collaborators and Other Affiliations template for one senior person — the 48-month co-author list from Crossref, project collaborators from NSF's own award record, written into the .xlsx Research.gov accepts, plus a separate summary of what was inferred, what was decided, and what was deliberately left blank for the PI to complete.

Heureka Labs · v1.1.0 1 file
nsf-current-and-pending grants

Build the SciENcv Current and Pending (Other) Support upload file for one senior person on an NSF proposal — current NSF awards read out of NSF's own record as PI and as co-PI, written in SciENcv's XML element order, with person-month years placed and effort, overlap statements and every non-NSF source left genuinely blank and named in a separate summary.

Heureka Labs · v1.0.0 1 file
nsf-synergistic-activities grants

Draft the one-page NSF Synergistic Activities document — five distinct examples chosen and compressed from a CV and NSF's own award record, rendered to a PDF whose page count, typeface, leading and margins are read back out of the file, with a separate summary naming what was selected, what was cut and why, and what only the senior person can supply.

Heureka Labs · v1.0.0 1 file
open-targets data

Ask the Open Targets Platform what is already known about a target or a disease — scored target-disease associations, the individual evidence records behind each score with their source and publication, tractability, and clinical candidates — over a public GraphQL API with no account. Resolves gene symbols and disease names to the Ensembl and MONDO identifiers the API demands.

Heureka Labs · v1.0.0 1 file
paperpush communication

Fill a journal or preprint submission portal from a manuscript directory — choose a venue, generate its submission file, extract the field values from the paper, validate them, then hand a signed-in browser to the author for the final submit.

Pachter Lab (adapted by Heureka Labs) · v1.3.0 1 file
polars-bio utility

Genomic interval operations and bioinformatics file I/O on Polars DataFrames — overlap, nearest, merge, coverage, complement, and subtract over BED/VCF/BAM/GFF/BigWig, plus streaming FastQC, with cloud-native paths.

K-Dense Inc. (adapted by Heureka Labs) · v1.3.0 8 files
pride data

Search the PRIDE Archive at EMBL-EBI for public mass-spectrometry proteomics datasets by keyword, organism, tissue, disease, instrument or software — read the free-text sample and data processing protocols, tell COMPLETE from PARTIAL submissions, and download raw, mzML, mzIdentML and search-engine output with checksums.

Heureka Labs · v1.1.0 1 file
proteinmpnn models

Design amino acid sequences for a fixed protein backbone with ProteinMPNN — inverse folding on CPU, chain and position constraints, homooligomer tying, and the soluble and CA-only checkpoints. Covers what native-sequence recovery does and does not tell you, and the silent ways a backbone stops being designed.

Heureka Labs · v1.0.0 1 file
protenix models

Predict all-atom structures of proteins, nucleic acids, ligands and ions with Protenix, an open AlphaFold3 reproduction — the input JSON grammar, MSA and template handling, checkpoint selection, and how to read pLDDT, PAE, pTM and ipTM without over-reading them. Running the model needs an NVIDIA GPU.

Heureka Labs · v1.0.0 1 file
protrek models

Search proteins by plain-language function, amino-acid sequence or 3Di structure with the ProTrek trimodal model. Covers the hosted search API that needs no account, which database and output-modality pairs actually return anything, similarity scoring with a negative control, structure queries with no local foldseek, and the CPU-local route with the 35M checkpoint and the faiss indexes.

Heureka Labs · v1.0.0 1 file
pxdesign models

Design de novo protein binders against a target structure with PXDesign — writing the target YAML in the residue numbering it actually reads, running the diffusion generator behind its AF2-IG and Protenix filters, and reading summary.csv without mistaking which PAE column is in angstrom or which ipTM is the interface. The generator needs an NVIDIA GPU; the checks in this skill do not.

Heureka Labs · v1.0.0 1 file
reviewer-scouting utility

Propose suggested reviewers for a manuscript submission, grouped so they cover every angle of the paper — find researchers who both publish on an angle and have recently published in the target journal, screen them against the author list for conflicts, and recover each one's published email address.

Heureka Labs · v1.0.0 1 file
rosetta-foundry models

Design proteins with RFdiffusion3 and predict structures with RoseTTAFold3 from the Rosetta Commons Foundry package — install, the checkpoint registry, the two incompatible input grammars, shipped defaults that contradict the documentation, and how to read RF3 confidence including a field that does not hold what its name says. Running either model needs a GPU.

Heureka Labs · v1.0.0 1 file
sam2 models

Segment and track objects with Meta's SAM 2.1 — point, box and mask prompts on a single image, automatic mask generation over a whole field, and propagation of a segmentation through a time-lapse. Covers where a natural-image model misreads microscopy, and what to check before trusting a count.

Heureka Labs · v1.0.0 1 file
scientific-critical-thinking utility

Evaluate scientific claims and evidence quality — assess experimental design validity, identify bias and confounding, and apply grading frameworks such as GRADE and Cochrane Risk of Bias.

K-Dense Inc. (adapted by Heureka Labs) · v1.1.0 8 files
sennet data

Query the SenNet cellular senescence atlas — senescent-cell datasets across 17 organs in human and mouse, the donors behind them with ages spanning the whole lifespan, and the tissue samples and files derived from each. Covers the Elasticsearch grammar the portal exposes, the source-sample-dataset hierarchy, which tier a dataset sits in, and how to download public files with no account.

Heureka Labs · v1.0.0 1 file
smaht data

Search and download SMaHT somatic mosaicism data — somatic variant calls, WGS and RNA-seq reads, and the donors, tissues and assays behind them. Walks donor to tissue to library to file and checks each file's tier before transferring. Cell-line benchmarking data is open with no account; every file from human tissue is protected under dbGaP, a route this skill describes rather than opens.

Heureka Labs · v1.0.0 1 file
spatial-phenotyping analysis

Spatial statistics on a segmented multiplex dataset — a cell-by-marker AnnData with coordinates, from CODEX, MxIF, IMC or MIBI. Phenotyping, neighborhood enrichment, cellular neighborhoods, Ripley's L and co-occurrence, with the per-slide graph construction and per-patient aggregation that keep the statistics honest.

Heureka Labs · v1.0.0 1 file
survival-analysis analysis

Time-to-event analysis with lifelines — Kaplan-Meier curves, log-rank comparisons and Cox regression on right-censored data. Tests the proportional-hazards assumption rather than assuming it, and covers stratification and RMST for when it fails.

Heureka Labs · v1.0.0 1 file
synapse data

Find and retrieve consortium datasets from Synapse (Sage Bionetworks) — resolve a syn id to entity metadata and annotations, walk a project's folders, and read an entity's access tier before attempting a download. Covers AMP-AD, the AD Knowledge Portal, PsychENCODE and HTAN. Most Synapse files need a free account; some need an approved data use certificate.

Heureka Labs · v1.1.0 1 file
uk-biobank data

Search UK Biobank's openly published field catalogue — 11,821 variables in 410 categories with participant counts, units, the instanced/arrayed shape that decides how many columns a field becomes, and the 172 retired fields the website hides but the download still carries — and report what an access application requires, applications being paused as of August 2026. Never fetches participant data.

Heureka Labs · v1.1.0 1 file
zenodo data

Retrieve data deposited alongside a paper on Zenodo — resolve a DOI or record id, read the record's licence and access status before any bytes move, refuse non-commercial or unstated terms rather than warn, pin a concept DOI to the version the paper actually used, and download selectively against a size budget.

Heureka Labs · v1.0.0 1 file