I wanted a fast, local way to query the small-body solar system (near-Earth objects, asteroids, close approaches) without hammering NASA/JPL APIs every time. So I built SpaceDB.

What it is

SpaceDB ingests the public JPL small-body and close-approach datasets into a queryable database, backed by ~1.5M object records, with a spacekit-powered 3D viewer for orbits.

The data pipeline

Four feeds go in, and they do not agree with each other about anything.

  • SBDB is the catalog: one row per object with its osculating elements at an epoch.
  • CAD is the close-approach table: one row per encounter, so an object appears many times.
  • Sentry is the impact-risk listing, with cumulative probabilities and Palermo scale values.
  • NHATS is the accessible-target list, keyed on mission delta-v rather than on anything astronomical.

The normalizing pass is most of the work. Designations are the first problem: the same rock is 433, 433 Eros, A898 PA and 1898 DQ depending on which feed you asked, so everything is resolved to the SPK-ID and the human-readable names are kept as an alias table rather than a key. Units are the second: distances arrive in astronomical units in one feed and lunar distances in another, diameters in kilometres where they exist and nowhere at all where they don't, and epochs as Julian dates that need converting before a timestamp column will accept them. Absent is not the same as zero, so every one of those columns is nullable and the ingest refuses to invent a value.

After that it's ordinary queries:

GET /api/objects?neo=true&diameter_min=0.5&sort=close_approach_date

That is, "near-Earth objects bigger than 500m, soonest approach first."

The indexing is deliberately boring. Partial indexes on the neo and pha flags, since those subsets are tiny compared to the main-belt bulk that makes up most of the row count, and a composite index on object plus approach date for the close-approach joins. Ingests are idempotent and keyed on the SPK-ID so a re-run patches rather than duplicates, which matters because JPL revises elements as new observations arrive.

Why bother

Once the data is local and indexed, exploratory questions are instant, and it pairs nicely with the AstroNN experiments. The other reason is politeness: the JPL APIs are a public good, and pulling the bulk files once a day is better behaviour than firing a thousand requests at them because I wanted to sort a list.

Source is on GitLab.