| Deployment | |
| Activity | |
| Python Versions | |
| Supported Systems | |
| Project Status | |
| Build Status | |
| Linting | |
| Code Coverage | |
| Code Quality | |
| License |
| Citation | | Documentation
RadDB archives xarray DataTree radar volumes as compact Parquet files and gives you a small, fluent interface to load, filter, crop, extract cross-section and plot them. It is network-agnostic: any DataTree with the standard xradar coordinate layout (NEXRAD, ODIM, IRIS, …) can be archived and analysed.
A radar is stored as a static LUT (per-gate geometry, generated once) plus one
POL parquet per volume (the variables), linked by an integer gate_id:
{archive_dir}/{radar}/LUT/{radar}_LUT.parquet # gate_id, lat/lon/alt, x_<epsg>/y_<epsg>, sweep, …
{archive_dir}/{radar}/{YYYY}/{MM}/{DD}/{radar}_{YYYYMMDD}_{HHMMSS}_POL.parquet
No-echo gates are dropped at archive time (default DBZH > 0), so the archive
stays small: a 12-sweep WSR-88D volume of 8,791,200 polar gates becomes 8.2 MB.
pip install raddb
pip install "raddb[viz]" # interactive Jupyter map and cartopy basemapsCore runtime dependencies: numpy, pandas, polars, geopandas, shapely, pyproj, xarray, pyyaml, pyarrow, matplotlib, netcdf4, zarr.
RadDB is a single, dual-role class:
- archive-bound —
db = RadDB(archive_dir, crs=…); use it toarchive()andopen(). - data-carrying — the object returned by
open()(call itrdf). It holds the data as a polars DataFrame (rdf.data) and exposes the query / convert / crop / cross-section / plot methods. Each of those returns a newRadDB, so calls chain fluently.
import raddb
db = raddb.RadDB(archive_dir="/data/raddb", crs=32614)A projected CRS is mandatory to write an archive and never needed to read one.
There is no default: the wrong projection is silently wrong.
raddb.lut.suggest_crs(longitude, latitude) tells you which one to pass.
# From saved DataTree files on disk (.zarr / .nc); the LUT is auto-generated:
db.archive(datatree_dir="/data/NEXRAD_datatree") # radar inferred per file
# ...or archive in-memory DataTrees directly:
db.archive(datatree=dt, radar="KTLX")
db.archive(datatree=[dt1, dt2], radar="KTLX")
db.archive(datatree={"KTLX": [dt1], "KMLB": [dt2]}) # multi-radarPass filter= to decide which gates ever reach the disk — the main control on
archive size:
db.archive(
datatree=dt, radar="KTLX", filter={"var": "DBZH", "logic": ">", "threshold": 20}
)rdf = db.open(time_period=("2024-06-12", "2024-06-13"), radars="KTLX")
print(rdf) # summary: gates, radars, time range, columns
len(rdf), rdf.columns(), rdf.radars()
rdf.start_time(), rdf.end_time()
rdf.extent() # [xmin, xmax, ymin, ymax] in `crs`
rdf.geographic_extent() # [lon_min, lon_max, lat_min, lat_max]
rdf.crs(), rdf.geographic_crs()columns= and filters= are pushed down into the scan, so only the rows you
asked for are ever materialised.
filtered_rdf = rdf.filter({"var": "DBZH", "logic": ">", "threshold": 20})
filtered_rdf = rdf.filter(
[
{"var": "DBZH", "logic": ">", "threshold": 20},
{"var": "RHOHV", "logic": ">", "threshold": 0.9},
]
)
df = rdf.to_pandas(with_geometry=True) # pandas + gate coordinates
gdf = rdf.to_geopandas() # GeoDataFrame (with CRS)
dt = rdf.to_datatree() # back to xarrayFilters are {"var", "logic", "threshold"} dicts, where logic is one of
==, !=, >, >=, <, <=. crs is an EPSG int (e.g. 32614), a
CRS object, or None.
box = rdf.crop_by_bbox(extent=[636_504, 676_504, 3_891_333, 3_931_333])
poly = rdf.crop_by_polygone("catchment.geojson")
disc = rdf.crop_around_point((656_504, 3_911_333), distance=20_000) # metres
# rdf.interactive_crop() # draw an AOI on a Jupyter mapcs = rdf.extract_cross_section(p1=(626_504, 3_911_333), p2=(686_504, 3_911_333))rdf.plot_ppi(sweep=1, variable="DBZH", save="ppi.png")
rdf.plot_rhi(azimuth=270, variable="DBZH")
rdf.plot_cappi(altitude=3000, variable="DBZH")
cs.plot_cross_section(variable="DBZH", save="xsec.png")Each plot draws into one Axes and returns the matplotlib artist, so you compose
panels by passing ax=.
rdf.filter({"var": "DBZH", "logic": ">", "threshold": 20}).crop_by_bbox(
extent=rdf.extent()
).plot_ppi(variable="DBZH", save="ppi_plot_example.png")db.inventory() # radars, volume counts, time ranges, size
db.inventory(detailed=True) # + LUT info, stored variables, day-by-day counts
db.inventory(datatree_dir="/data/NEXRAD_datatree") # DataTree files not archived yetdb.list_radars() # radars present in the archive
db.get_lut("KTLX") # the static LUT (polars)
db.get_radar_info("KTLX") # site location / sweep geometry
db.add_lut_projection("KTLX", epsg=32614)raddb/
├── __init__.py
├── main.py # the RadDB class
├── io_core.py # DataTree <-> DataFrame <-> Parquet + archive backends
├── lut.py # LUT generation / geo projection
├── aoi.py # AOI / crop / cross-section geometry
├── discovery.py # find_datatree_files + filename-time parsing
├── helper.py # filters, radar-name normalisation, timers
└── viz/ # plot.py (PPI/RHI/CAPPI/cross-section), interactive.py
- Projected coordinates /
crs. Generating a LUT with a projection (e.g.crs=32614) and the projected accessors (extent,to_geopandas) usepyproj, which needs the PROJ database. APROJ_DATA/PROJ_LIBinherited from another environment (a conda base env, a system PROJ) points at a proj.db of the wrong PROJ version and makes every projection fail with "no database context specified" —import raddbdetects that and repoints PROJ at the running interpreter's ownshare/proj(seeraddb/_proj.py;raddb.PROJ_DATAreports what it changed,Nonewhen nothing had to be). - Zarr / NetCDF. Saved volumes are read back with
raddb.open_any_datatree(engine auto-detected); either format works.
See LICENSE (MIT).