How we research strains

Every strain page on The Bake Down follows the same editorial framework. This page explains which information is measured, which information is estimated, how search works, and what the numbers cannot tell you.

Where the data comes from

There are two clearly labelled data tiers on the site: lab-sourced figures and estimated figures.

Lab-sourced figures. For cultivars with matching published laboratory records in our Cannlytics dataset, the THC and terpene percentages shown as lab-sourced are calculated from those published results. These pages carry a green Lab-sourced figures badge and link to Cannlytics. They are measurements from published testing data, not a measurement of the particular product you may be holding.

Estimated figures. Most cultivars do not have a matching published laboratory record in our dataset. Their profiles are generated with a language model using a fixed output structure and editorial plausibility checks. Estimated THC and terpene figures are estimates; effects, flavor, aroma, genetics, growing information and other descriptive fields are editorial/model-generated. They should not be read as laboratory results or as claims about what every example of a strain contains.

Search results. Attribute searches do not ask the model to invent a list of strains. They rank the site's existing strain catalog using the attributes already recorded on those pages. A search for a strain name resolves against the catalog, including known aliases and close spelling matches. Only when a name is not in the catalog does the site generate a new estimated profile live.

When a requested strain is not in the catalog, the site generates one estimated profile and stores it in persistent server-side storage. Future searches for that normalized strain name return the saved profile rather than generating a new answer for each visitor. This keeps an estimated profile stable while still allowing it to be reviewed and updated through the site's editorial audit process.

Photography is licensed under Creative Commons and credited at the bottom of every page.

Why our numbers differ from other sites

Different sources can legitimately produce different numbers because cannabis potency varies by cultivar, grower, harvest, batch and testing laboratory, and because some consumer-facing figures represent a particular high-testing batch.

As an internal audit of the site's earlier estimated ranges, we compared the midpoint of each prior estimate with the matching Cannlytics total-THC value for 30 strains. The prior estimates were 6.5 percentage points higher on average (median 7.0 points). That comparison describes the earlier estimates; it is not a claim that every current estimate is exactly 6.5 points high.

How we report THC

On lab-sourced pages, the primary THC figure is the mean of the matching published test results in our dataset. It is an aggregate across the available records, not a test of a particular jar. The dataset does not expose a reliable sample count for every cultivar, so we do not imply that every average has the same statistical weight.

Some lab-sourced pages also show a reference range (estimated). This is an editorial estimate intended to provide context around the lab average. It is not a survey of dispensary menus or consumer websites and should not be read as a market-wide reported range.

On estimated pages, THC is presented as an editorial range informed by the cultivar's reported era and breeding context. The broad guide used in the system is:

These are calibration bands, not measurements and not guarantees about a cultivar. If potency matters to a purchase decision, the certificate of analysis (COA) for that specific product is the authoritative source.

Terpenes and effects

On lab-sourced pages, terpene percentages shown as lab-sourced come from the published laboratory dataset. On estimated pages, terpene amounts are editorial estimates. Terpenes are listed in order of the profile's estimated or measured prominence and use a consistent nine-terpene vocabulary across the site.

Effects are relative prominence scores, not percentages of people, probabilities, clinical measurements, or predictions. The site displays them as qualitative strengths such as Very prominent or Moderate. They describe the character commonly associated with a strain, not what an individual will necessarily feel.

“Commonly associated with” is descriptive language only. It is not a statement that cannabis treats or prevents a medical condition.

How search works

The site's attribute search ranks the existing 300-strain catalog; it does not generate six fictional matches. Matching considers the strain name and aliases, type, THC range, listed effects, terpenes, flavor and aroma notes, associated experiences, and lineage. Common-language synonyms are mapped to the site's controlled vocabulary so a query such as “sleepy and citrusy” can find relevant recorded attributes.

Results are returned only when the catalog contains a meaningful match. The result card identifies whether the strain's figures are lab-sourced or estimated. This prevents the search model from inventing a name, THC range, or explanation that is not represented in the site's data.

Searching for a strain name first checks exact names and aliases, then close spelling matches. If no catalog match exists, the site checks its persistent generated-profile store. Only if no saved profile exists does it generate a new estimated profile. That profile is validated, stored, and explicitly labelled Estimated figures.

How "more strains like this" works

The similar-strain suggestions on each page are computed from the data on the pages themselves, not from popularity or paid placement. You can switch between five definitions of similarity:

Where a strain has no meaningful match on a dimension, that option is hidden rather than filled with a weak suggestion.

How profiles stay current

Stable does not mean permanent. The site runs an automated weekly maintenance pass. Published laboratory data is checked against the current Cannlytics dataset, and material differences are flagged for review rather than silently changing a published page.

Estimated profiles are reviewed in a rotating batch. The weekly process works through the catalog from the oldest or least recently checked profiles so the estimated catalog receives a roughly quarterly review cycle as the site grows. A model review can flag a profile for human attention, but it does not itself establish factual truth and does not automatically rewrite the published profile.

This maintenance system is intended to catch stale or inconsistent information while preserving stable results for readers. A product-specific COA remains the authoritative source for what a particular product contains.

What this is not

This is a reference for understanding strain information. It is not medical advice, it is not a substitute for a product-specific lab certificate of analysis, and it cannot tell you what a particular jar on a particular shelf contains.

Cannabis affects people differently. Tolerance, format, dose, setting and individual chemistry all matter. Strain-level information is a guide to reported character, not a prediction of an individual's response.

How this site was built

The Bake Down was built with Claude, Anthropic's AI assistant. Claude wrote the code, generated the strain profiles marked as estimates, and built the tooling behind the site, including the search and similarity systems, the terpene reference and automated consistency checks.

Editorial decisions, verification and standards are human. Model output is treated as a draft and is subject to validation rather than being accepted as authoritative. During development, errors were found and corrected, including a badge/data-validation error that temporarily caused missing laboratory data to be treated incorrectly. That experience is one reason the current system separates source status from generated content and validates it explicitly.

The visual identity was developed with a human graphic designer, who shaped the bakery-case concept, palette and typography.

Corrections

If something here is wrong, we would rather know. Strain names overlap, lineages can be disputed, and the same name can describe different plants in different markets. Corrections and sourcing are welcome.