SoilTechSoilTech
Home/California ATBI

California All-Taxa Biodiversity Inventory

A statewide initiative to inventory all the species in the biodiversity hotspot of California. SoilTech is supporting this effort through cataloguing all of California's soil organisms.

Visit the initiative at calatbi.org →
1,100+
Soil samples processed
10,000+
Organisms extracted
Data access

Download the data

01 — Download

Summary export

A per-sample summary — organism and nematode counts and a diversity index for every soil sample — in a single CSV. The quickest way to get a feel for the data.

02 — Download from AWS

Full dataset on S3

The complete curated dataset — raw laser scans, high-resolution and processed imagery, and per-organism data (4+ TB) — lives in Amazon S3 (US East 1, requester-pays). See the access steps below.

03 — Cite

How to cite

Released under CC BY 4.0. Please cite the dataset when you use it.

SoilTech (2026). California All-Taxa Biodiversity Inventory: soil organism occurrence records. Version 2026.1. SOIL LLC. https://soiltech.bio/data. Licensed under CC BY 4.0.
LicenseCreative Commons Attribution 4.0 International (CC BY 4.0) — free to share and adapt with attribution.Read the terms →
Data access documentation

How to access the data on AWS

The full California ATBI dataset is hosted on Amazon S3, in the N. Virginia region (us-east-1). The data is completely open — there are no restrictions on how you use it, beyond the attribution asked for by its CC BY 4.0 license. The buckets are requester-pays, which means the data itself is free; you simply cover Amazon's standard charges for the requests and transfer when you download. A quick look costs a few cents, and the full multi-terabyte archive scales with how much you pull.

Access is granted per AWS account. You will need an AWS account of your own, then send us either your 12-digit AWS account number or the IAM user or role ARN you would like us to authorize. We add that identity to the bucket policy, and from there you can read and download the data directly — through the AWS Console, the AWS CLI (aws s3 sync … --request-payer requester), or any S3-compatible tool. To request access or ask a question, email team@soiltech.bio.

Before downloading several terabytes, you can see exactly what is inside. We publish an S3 Inventory of every object in the bucket — path, size, last-modified date, and checksum — and surface it in a searchable browser right here on the site. Use it to scope a download, confirm a sample is present, or just explore the archive.

Building on the California ATBI? We'd like to hear about it.