Runoff and green infrastructure reduction, for small sites

The curve number is not the hard part.

Defending it is. You draw a boundary, and CurveNumber reads the soils, the land cover, the slope and the design rainfall out of the public datasets, crosses them, and computes the reduction between the existing and the proposed condition. Then it prints, beside every figure, either a citation with an edition, an assumption with what it is worth in curve number units, or the name of the person who changed it and the reason they gave. Where nothing can honestly be derived, it says so and withholds the number rather than filling the gap.

It is built for the civil engineer or site planner who will seal the drawing and then sit across from a reviewer who is paid to doubt it.

Two sites free, then USD 100 per site, charged the first time a site is computed. No subscription. The quick check needs no account at all.

The head of a runoff reduction report: report identifier CNR-3DB391A663FB, run identifier, generation time and engine version, then the site name and a headline figure of 1,478 cubic feet of reduction with existing and proposed depths and volumes beneath it. Below that, a section explaining where the water went, and the retained volume reported on its own.
The deliverable. Head of the example report in the repository, rendered by app/src/cnapp/report.py. The identifier is a hash of the run and excludes the generation time, so two printouts of the same run carry the same identifier and a reviewer holding both can tell at a glance.

What is actually being sold

Every quantity is one of three things, and it says which

Nothing in this engine is a bare number. A quantity carries its magnitude, its unit and how it came to be, and construction fails if that is missing. A reviewer's question is always some version of "where did this come from", so the answer is attached to the figure rather than reconstructed afterwards.

derived

Retrieved from a published source

The citation carries an edition, because TR-55 1986 and TR-55 1975 disagree and a citation without an edition cannot be checked. A derived value with no edition on its source raises rather than saving.

USDA NRCS TR-55, 1986, 2nd ed.,
Tables 2-2a to 2-2d
defaulted

An assumption, priced in the answer

Each default names a registry entry, and the entry states what the assumption is worth in curve number units or in inches of runoff, and what evidence would displace it. Adjectives are not accepted there.

condition.default = Fair
worth 5 to 12 CN units
overridden

Changed by a person, with a reason

Who, when, the prior value and the reason given, written to an append-only record that cannot be edited afterwards. The published value stays visible next to the attested one, because the arrow between them is the record.

no group → C
soil boring log, 2 borings, tested

An override is not a text box

The interface states what the choice is worth before it is made, asks for a reason in proportion to that figure, and names the account the sentence will be filed against. An engineer who has to write "because the software said B" tends to go and look instead, which is the entire intent.

The soil survey rates no hydrologic group for a large share of urban land, which is the ordinary case on the sites this product is for. That is not a blank to fill with B. It is a question, and it goes to the person who can answer it.

An answered question in the interface: map unit Urban land, marked overridden. The survey published no group; a person attested group C. The share of site is 97.8 percent, and the reason reads: soil boring log, borings B-3 and B-7 logged 1.2 metres of silty fill over Codorus alluvium, infiltration tested at 0.3 inches per hour, which is group C.
An answered question, which is not a resolved one. Both halves stay on the record.

The part that is unusual

It refuses, in writing, rather than filling the gap

Any tool can put a number in every cell. The failure mode that costs an engineer their afternoon in a review is the number that looks entirely reasonable and is quietly wrong, and it is nearly always produced by something being filled in silently: an unrated soil map unit read as B, a pond given CN 98, a dual soil group collapsed to one letter.

So the engine separates the states that have different remedies and keeps them separate all the way to the screen. A service being down and a site being outside the survey are not the same problem and are never the same message. Where a material share of a site is unresolved, no headline number is produced at all.

The threshold is a tenth of the site, and it is recorded as a judgement rather than a finding. At a twentieth unresolved, the compounded error on a site whose gap is really group D is already 10.6 percent of the design volume, so a tenth is not the point where the gap becomes harmless. It is the point where warning stops being a proportionate response.

segmentation.py, MATERIAL_GAP_FRACTION = 0.10, with the arithmetic in the comment above it

The derived data panel refusing to answer. A red banner reads: no curve number for this site yet. No part of this 4.000 acre site has a resolved hydrologic soil group, so there is no cell of TR-55 Table 2-2a to read and no curve number to report. Below it, two open questions, and a soil table showing map units marked not populated.
A real site, refused. The Fairfax test area is 97.8 percent Urban land with no published hydrologic group. The area is neither defaulted nor dropped: it is carried with its acreage and its map unit name, and it becomes a question.

What a refusal reads like

NLCD class 11, open water, refused

TR-55 has no open water row and open water is not a rainfall to runoff transformation. A pond is a storage element: what leaves it during a storm is set by its stage, its surface area and its outlet, none of which a curve number represents. The common assignment of 98 or 100 asserts that the whole surface rainfall leaves instantly, which on a site otherwise at CN 69 on group B with a pond over a tenth of the area moves the answer from 0.670 to 0.880 in at P = 3 in, and from 0.019 to 0.116 in at P = 1.2 in.

Supply instead: the water body as a routing element with a stage, area and discharge relationship, or its area excluded from the curve number computation with the exclusion stated in the report.

crosswalk.py, REFUSED[11], quoted in full. Eight land cover classes are refused this way, each with its own reason and its own remedy.

The defaults registry

An assumption that states its own price

Ninety-eight values in this engine are defaults rather than anything read off your site. Every one of them is written down in a registry, classified by what would have to happen for it to change, and carries two fields that turn a lookup into guidance: what it is worth in the answer, in computed figures rather than adjectives, and what evidence would displace it.

Fifty-four of them say no external source exists on their face, because nobody publishes the figure at all. Seven more say the source exists and was not checked, which is a different admission with a different remedy, and the two are never merged. The registry is meant to be published free and citable.

That is also how the interface can be quiet. A user shown thirty defaults confirms none of them, so the engine ranks its own assumptions by what they are worth on this site and surfaces the two or three that move the answer.

A section of the report headed: the assumptions that carry this answer. Three assumptions are listed, each marked defaulted with its registry key, class and confidence, then a line beginning worth in the answer, then a line beginning displaced by. The dual hydrologic soil group entry reads: enormous, open space in good condition on A over D is CN 39 drained against CN 80 undrained.
The same registry, on the face of the report. Six ranked assumptions reach the page; the rest are counted and pointed at rather than printed.
3.42×

Runoff at P = 3 in between reading a B/D soil as drained and as undrained. 0.365 in against 1.250 in. The engine reports B/D as B/D and records the resolution as its own assumption, because collapsing the two would make "the map says D" indistinguishable from "the map says B/D and we assumed no drainage".

61%

How high a composite curve number runs on NLCD class 22 if the TR-55 residential row is used as the pervious cover and the measured impervious share is added on top of it. That row already contains 38 percent impervious, so the share is counted twice. Nothing about the answer looks wrong.

0.041in

Initial abstraction at CN 98, against 1.279 in at CN 61. Which is why the depth criterion in this engine is the storm over the segment's own initial abstraction, and not an absolute number of inches, which would be wrong in both directions at once.

registry.py and crosswalk.py. Every runoff figure in those files is computed with cn.runoff_depth at lambda = 0.20 on the tabulated curve numbers, and pinned against the code in tests/test_registry_arithmetic.py

Where the inputs come from

Public data, named, dated and cached

You supply the boundary and the design storm. Everything below is retrieved, and each retrieval is recorded with its URL, its status, the timestamp and a hash of the response body, so a figure in a report traces to a specific server response on a specific day.

The boundary is drawn on the map, or edited vertex by vertex, or cut with a hole where a building is excluded. The enclosed area is computed on the ellipsoid and shown while you draw, because an acreage that appears only after you commit is an acreage nobody checks.

The drawing surface: a topographic basemap with a four acre rectangular site boundary drawn on it in teal, an address search box and drawing tools floating over the map, a legend listing the drawn area, and a large readout showing 4.000 acres and 174,241 square feet of enclosed area.
The drawing surface. Basemaps are USGS topographic and imagery, and the tool says so on the map, including when you have enlarged past the resolution the source actually has.
Soil map units and hydrologic groups USDA Soil Data Access, SSURGO
Map unit names, keys and published groups, with dual groups kept as B/D and unrated units carried as named gaps rather than filled.
Land cover and measured imperviousness Annual NLCD, via USGS and MRLC
Class shares and a measured impervious fraction from the fractional raster. The impervious share is never inferred from the class code, because the published composite rows already contain one.
SlopeUSGS 3DEP elevation
Computed server side over the boundary and reported in percent rather than in degrees, because the two are not interchangeable and confusing them understates slope. It feeds time of concentration, which also needs a flow length, and no elevation raster contains one, so that stays your input.
Design rainfallNOAA Atlas 14, PFDS
Depth and duration for the point, with the edition recorded. A design storm is a depth and a shape, and the shape is stated separately.
Address lookupUS Census Geocoder
Used only to place the map view. The returned point is interpolated along a street centreline and offset to one side, so it is never used as an area of interest, and the interface shows the block range it came from.

engine/src/curvenumber/sources/. These are unfunded public services, so every response is cached to disk by request, and the test suites run offline against a recorded fixture set rather than reaching them.

Read this before you buy anything

Where this actually is

This is pre-launch software. It has never had a paying customer, and the version number is 0.8.0. If a page like this one is going to argue that provenance is the product, it has to be candid about its own, so here is what is not finished.

The transcribed tables have not been independently checked

The published curve number tables and the rainfall distributions in this engine were transcribed by one person and checked by nobody. The code records that as CHECKED_BY = None and a test asserts it stays honest, so every report that uses them prints "checked by NOBODY" beside the citation rather than leaving a reviewer to assume otherwise.

Two specifics. The cover table shipped is a 22 row seed subset of TR-55 Tables 2-2a to 2-2d, not the whole of them. And one row of the antecedent runoff condition table is under active suspicion: its local gradient does not match its neighbours in a way no smooth published relation would produce. It is left exactly as transcribed, with the doubt recorded in an open test, because the remedy for a transcription doubt is a person with the published document in front of them and not a guess from here.

engine/src/curvenumber/tables.py and distributions.py, and tests/test_published.py::TestTheArcTableIsAnOpenItem

An unchecked transcription is a real defect, and the reason to say so on the front page rather than in a footnote is that it is exactly the sort of thing that gets discovered by a reviewer instead. Every figure this product prints requires independent verification in any case. What the product is for is making that verification cheap: the citation, the edition and the assumption are already next to the number.

Also not finished

  • Peak discharge is not implemented. Before it is, the lag to time of concentration relation has to be verified against the primary source. The engine computes volumes and depths, routes a hydrograph through a practice, and does not give you a peak flow for a pipe sizing.
  • The coverage is the conterminous United States, because the datasets are. Land cover classes that only occur in Alaska are refused by name rather than mapped to something that looks plausible.
  • One metre land cover is resolved but not read. The Chesapeake Conservancy assets are located and identified; each is a whole county at 143 MB and reading them needs a range reader that is not written.
  • Payment is not wired. There is no checkout yet. How you pay today is on the pricing page, and it involves a human being.
  • The assistant ships switched off. There is an optional assistant that answers questions about a site you have computed, and it is off by default on every deployment. It may never produce a number: every figure it states has to be one the record already holds, every citation has to be something it retrieved in the same turn, and a response that breaks either rule is blocked before you see it, by a deterministic check rather than by asking the model nicely. What it has not had is the expert review panel that would let anyone claim it is correct, and its own evaluation report says so in its own words. Where it is switched on it sends your site's record to a commercial model vendor, which the privacy page describes item by item.

README.md, "Honest gaps", which the codebase keeps current

Price

Two sites free, then USD 100 per site

A site is one boundary with its existing and proposed conditions, and it is charged the first time it is computed rather than when it is created, so you can set one up, look at what the public data says about it, change your mind and pay nothing. Deriving the public data is free and does not consume anything. There is no subscription, no seat count and no annual commitment.

The stateless quick check is free and needs no account. Nothing about it is stored, which also means it carries no audit trail, and the audit trail is the part a reviewer relies on.

What you get for the hundred dollars

A report, in HTML and as a JSON companion generated from the same build so the two cannot disagree, holding the reduction and its composition, the segment table with the citation and origin on every value, the ranked assumptions with their worth and what would displace them, every departure from a default with who made it and why, the state of each public service at the time of the run, and the identifiers needed to reproduce it.

Retention and filtration are reported separately and are never added together anywhere in it, because many jurisdictions credit the first and not the second.