Kinoteca user guide

Kinoteca is a platform for exploring and downloading kinase inhibitory-activity data. It brings together kinase-target metadata, compound structures/properties, and measured inhibitory activities in one place so you can browse, filter, visualize, and export exactly the slice of data you need — without writing any code or SQL.

This guide covers what the database is useful for, how to run a filter, how to read and export the results, and how to save a filter for reuse.

1. What's in the database, and what it's for

Kinoteca is a data-exploration and curation tool, not a predictive-modeling tool — it won't train or serve models for you, but it will help you quickly answer questions like:

  • "Which compounds have been tested against this kinase, and how active were they?"
  • "Which kinases in this pathway/family have compounds with a molecular weight under 500 and good drug-like properties?"
  • "What raw, pre-fusion measurements support this curated activity value?"

Three kinds of information are linked together:

  • Kinases — identity (name, gene, UniProt ID), kinome classification, structural links (PDB IDs, an AlphaFold model), and group/pathway memberships.
  • Compounds — structure (SMILES, InChI key) and physicochemical properties (molecular weight, ALogP, H-bond donors/acceptors, polar surface area).
  • Activities — measured inhibitory activity of a compound against a kinase. Each curated activity value is a fusion of one or more individual ("raw") measurements; a data-quality metric called MUE (a dispersion measure across the replicate measurements that were fused together) tells you how consistent those underlying measurements were. Data that failed an acceptance threshold for MUE upstream never became a curated value in the first place, so every curated activity you see has already passed that quality bar.

You can browse kinases directly (Kinases in the navigation bar) to look one up by name, gene, or UniProt ID and see its classification, activity summary, group tags, and 3D structure. But the main way to work with the data is the filter pipeline, described next.

2. Running a filter

Running a filter requires being logged in (see Create account in the navigation bar if you don't have one yet — it's free and immediate, no approval step). The filter pipeline is a three-stage wizard, reached from Filter in the navigation bar:

Stage 1 — Choose kinases

Search by name/gene/UniProt ID and/or narrow by facet (kinome classification, pathway, target group, and any other curated groupings). You can then either:

  • select one or more individual kinase cards and click Continue to carry exactly those kinases forward, or
  • click Continue with nothing selected to carry forward the entire set of kinases matching your current search/facets.

Stage 2 — Molecule properties and MUE

Narrow the compounds by physicochemical property (molecular weight, ALogP, H-bond donors/ acceptors, polar surface area) using range sliders overlaid on a live histogram of the selected kinases' data, plus a maximum MUE dispersion threshold to exclude noisier curated activities. The match counts update live as you adjust a slider. Click Continue when you're happy with the ranges (or leave everything at its full range to skip this narrowing).

Stage 3 — Choose datasets

A "dataset" here means one (kinase, activity type) pair — e.g. "EGFR, IC50". Pick which of the resulting datasets to keep, optionally first narrowing the list to datasets with at least a minimum number of distinct compounds. Confirming this stage creates your result set.

3. Viewing your results

Once the wizard finishes, you land on the result page for that filter, which shows:

  • summary counts (kinases, datasets, compounds, activities) and a per-kinase breakdown table;
  • a plain table of the matching activities (capped to the first 500 rows for readability — the full set is still available via export, see below);
  • a histogram of the activity distribution, plus one histogram per physicochemical property, computed over the filter's unique compounds;
  • structural clustering — compounds grouped by chemical similarity (Tanimoto distance over Morgan fingerprints), largest cluster first. This step is skipped above a few hundred unique compounds, since it gets slow at that scale — narrow your filter further if you need it;
  • scaffold grouping — compounds grouped by their Bemis-Murcko core scaffold, largest group first. Unlike clustering, this has no size cap.

4. Exporting your data

The result page offers three downloads, each covering exactly the activities in that filter set:

Export Format Contents
Curated CSV CSV One row per curated (fused) activity: compound identifiers (ChEMBL ID, InChI key, SMILES), kinase name/UniProt ID, activity type, the relation operator (=, >, <=, ...) and value/units as measured, the MUE dispersion value, and how many raw records were fused into it.
Raw CSV CSV The semi-raw, pre-fusion data underlying the same curated activities: one row per individual measurement, with its own relation/value/units, its original source and external ID/reference, and its upstream data source (e.g. a ChEMBL assay ID) for full traceability back to the original measurement.
Compound SDF SDF (structure file) One structure block per unique compound in the filter set, tagged with its identifiers, physicochemical properties, and a summary of every curated activity for that compound within this filter set — so the bioactivity data travels with the structure into any cheminformatics tool that reads SDF.

Use the curated CSV for the standard "one number per compound/kinase/type" view, the raw CSV when you need to see or audit the individual measurements behind a curated value, and the compound SDF when you want to take the structures (with their data attached) into another tool.

5. Saving and reusing a filter

From a result page, give it a name and click Save filter to store the pipeline definition (the kinases, property ranges, and dataset selection you chose — not a snapshot of the results themselves). Saved filters appear under Saved filters in the navigation bar, where you can:

  • Run a saved filter to re-execute it against whatever data exists right now, producing a fresh result set — this is the main point of saving a filter: as the database grows, re-running a saved filter picks up new matching data automatically, rather than being frozen at the moment you first ran it;
  • Save again under the same name to overwrite that filter's definition with your current selections (there's no separate rename/edit step — saving again is editing);
  • Delete a saved filter you no longer need (this only removes the saved definition, not any result sets it previously produced).