Commercial road-surface intelligence · North America + Europe

96.09% of roads across North America and Europe classified as paved or unpaved.

License confidence-scored paved or unpaved road-surface classifications. Sherpa classifies all 66.40 million production road sections using multispectral imagery and broad geographic context. With a commitment to flexibility, data delivery is matched to your requirements.

Road Surface Classification is the lead product in a broader road-data portfolio, alongside Scenic Road Score and Modeled Traffic, a consistent comparative network-use metric.

Partner brief

Validated performance: 94.4% and 86.9% accuracy across 1.95 million imagery predictions.

ROC-AUC measures how reliably the imagery model ranks paved examples above unpaved examples across every possible decision threshold. A score of 50% would be random ranking; Sherpa achieved 94.4%. Out-of-fold means each benchmark prediction came from a model that was not trained on that road's local geographic cell. This makes the test harder to pass through geographic memorization.
Live continental data Surface intelligence
Full atlas
Surface-classified roads over satellite imagery
Tan indicates higher unpaved confidence; dark indicates higher paved confidence. Choose a close-up to inspect road detail.
Higher unpaved confidenceHigher paved confidence
Switch between the three finished layers, then open the full atlas for basemaps, opacity controls, and continental exploration. Drag with the left or middle mouse button; use +/− to zoom.
01

Lead data product

Road Surface: Paved or Unpaved

Give routing, mapping, fitness, mobility, and outdoor products one consistent surface answer across the full production network. Sherpa classifies roads with missing tags and reclassifies tagged roads using newer imagery and geographic context.

Explore surface data

The commercial gap

OSM supplied usable surface tags for 37.8%. Sherpa independently classifies the full network.

Across the same 66.40 million North American and European production road sections, 25.08 million carried a usable OSM paved or unpaved source tag. Sherpa does not preserve those tags as final answers. It reclassifies all 66.40 million sections using newer imagery, road and network context, language, land use, structures, terrain, development, and modeled road importance. Of those roads, 41.32 million, or 62.2%, had no usable OSM surface tag at all.

Same production network, same denominator

OSM source coverage compared with Sherpa classification coverage

This measures whether a road had usable source metadata before Sherpa. OSM tags support training and evaluation where present, but the finished Sherpa layer reclassifies those roads too. It does not copy 37.8% and fill only the remainder.

OpenStreetMap source tags25.08 million of 66.40 million sections
37.8%
Sherpa finished surface layer66.40 million of 66.40 million production sections
100%
Same footprint, two context viewsThe low-contrast OSM map makes the coverage gap easiest to read. Switch to satellite context to inspect the same evidence over processed Sentinel-2 imagery.
Usable OSM source tags4,486 of 82,342 road sections
Road sections with usable OpenStreetMap paved or unpaved source tags in a selected Louisville-area footprint
OpenStreetMap source coverage: only the blue and ochre roads carried a usable paved or unpaved tag. In this deliberately selected sparse case, that is 4,486 of 82,342 road sections, or 5.45%. The measured North America and Europe average is 37.8%. Map and road data © OpenStreetMap contributors.
Sherpa finished surface layer82,342 of 82,342 road sections
Every road section in Sherpa's selected Louisville-area production footprint classified as paved or unpaved
Higher unpaved confidenceHigher paved confidence
The exact same footprint, independently reclassified: Sherpa assigns a paved or unpaved classification and confidence to every production road section, including roads that began with an OSM surface tag. The result is ready for route choice, warnings, map styling, coverage analysis, and QA. Map and road data © OpenStreetMap contributors; context data: Overture Maps.
46.45Mroad sections classified as paved
19.96Mroad sections classified as unpaved
41.32Mclassified with no usable OSM surface tag available
100%independently classified by the Sherpa system

The engineering beneath one road color

The visible answer is only the tip.

One paved or unpaved value can draw on multispectral imagery, language, road geometry, graph structure, terrain, land context, buildings, and billions of modeled journeys. No single input decides the answer. The depth comes from independent measurements that reinforce or correct one another.

A Cycles-rendered iceberg with a small exposed crown, a physical water surface, and a much larger illuminated mass below
Delivered road intelligencePaved or unpavedConfidence and pre-reclassification source-tag status included

Structured fusion

8,401 active feature columnsacross two production matrices that combine numerical, categorical, image, text, graph, and geographic context

Specialized model stages

Spatial vision modelsConvolutional and transformer candidates trained with geographic folds
Learned representationsVision-language, self-supervised image, and compressed road-name context
Road-level reconciliationMultiple chips become probabilities, uncertainty, spread, extremes, entropy, and support

Engineered measurements

115.71M road-centered chipsRGB, red, NIR, NDVI, reflectance, masks, gradients, and along-road statistics
Road and graph contextClass, access, lanes, speed, geometry, topology, neighbors, and network density
Road-name meaningNormalized names encoded from 768 dimensions into 64 cached numerical dimensions
Surrounding geographyBuildings, land use, crops, forests, water, soils, terrain, climate, places, and development
5.36B contributing pathsRoad importance derived from 15.05 billion VIIRS-guided routing attempts

Continental source systems

Sentinel-2RGB, B04 red, B08 near-infrared, spectral indices, masks, and composites
OpenStreetMapKnown examples, names, geometry, topology, access, road classes, and network relationships
Overture MapsTransportation, buildings, land use, land cover, water, places, addresses, and divisions
VIIRS and routingNighttime-light activity centers and a purpose-built high-throughput C++ routing system
Physical environmentElevation, terrain, water, vegetation, soil, climate, agriculture, and regional context
23.22 trillionfour-channel image values processed
300.54 billionstructured feature-value slots across the five-region production matrices
15.05 billionsampled origin-destination routing attempts
69.10 millionclassified road sections in the five-region finished release
01Training examplesNormalize mapped surface labels and assign spatial folds for held-out reclassification.
02Road-centered imageryAlign RGB, red, and NIR; mask the road; compute spectral context.
03Spatial vision modelsGenerate held-out vision predictions for tagged roads and full-model vision predictions for untagged roads.
04Road-level aggregationCombine multiple chip predictions into mean, spread, extremes, entropy, and support.
05Structured fusionFuse context into held-out final classes for tagged roads and full-model classes for untagged roads.
06Delivery and reviewExport class, confidence, pre-reclassification source-tag status, and partner-matched identifiers.
Technical detailHow the imagery, model stages, text, and structured fusion work

Imagery preparation

Sentinel-2 Global Mosaic RGB, B04 red, and B08 near-infrared are sampled around roads. B04 is aligned to the NIR reference grid where needed. The chipper applies a road mask and nodata/validity filters, then measures center-road and along-road red, NIR, NDVI, DVI, simple ratio, brightness, RGB ratios, saturation, excess green, quantiles, threshold shares, and gradients.

Vision sequence

Five geographically grouped folds produce untouched stage-one predictions. Production model stages aggregate per-chip probabilities, uncertainty measurements, and learned representations into road-level evidence. The finished feature stacks combine convolutional predictions with vision-language and self-supervised representations where available; these become inputs to the structured model rather than a vote-only ensemble.

Road-name encoding

Primary, alternate, official, local, English, address-street, and reference fields are normalized and deduplicated. A pretrained sentence encoder creates 768-dimensional representations, reduced and cached as 64 numerical dimensions so names can inform the model without millions of sparse categories.

Final fusion

CatBoost combines numerical, categorical, image, text, graph, and geographic representations into the finished road-level decision. Each delivered classification includes the Sherpa class, confidence score, and whether a usable OSM source tag existed before reclassification. Tagged roads receive spatially out-of-fold predictions; untagged roads receive full-model inference. Counts use one mapped road section per distinct OSM way identifier, not one unique named road or routing graph edge.

Completed surface benchmark

Proven road surface classification.

Evidence: 1,947,588 balanced road predictions across five . Roads in the same H3 resolution-6 cell remain in the same fold, so the same local cell cannot appear in both training and validation.

94.37%
94.43%
86.86%accuracy
86.86%
PR-AUC summarizes the balance between finding paved examples and avoiding false paved classifications across possible thresholds. Higher is better, especially when class frequencies vary. Macro F1 gives paved and unpaved equal importance, then combines precision and recall for both classes. It prevents a larger class from dominating the score. Spatial folds divide the benchmark by geographic cells. Every road in one local cell stays on the same side of the train and validation split, preventing the same place from appearing on both sides of the benchmark.

Scope: this benchmark measures the spatially held-out imagery stage on its own. The finished production classification adds the structured geographic, graph, language, and routing context shown above.

Confidence-ranked coverage

Higher-confidence imagery decisions are substantially more accurate.

Selective results from the same untouched benchmark. is distance from the 0.5 decision boundary, not a claim of calibrated probability.

Here, confidence measures how far the imagery model's paved probability is from an even 50-50 decision. It ranks decisiveness but is not presented as a perfectly calibrated real-world probability.
All predictions100.00% coverage
86.86% accuracy
Confidence ≥ 0.674.90% coverage
94.01% accuracy
Confidence ≥ 0.768.63% coverage
95.28% accuracy
Confidence ≥ 0.860.76% coverage
96.50% accuracy
Confidence ≥ 0.948.14% coverage
97.90% accuracy
Coverage retainedAccuracy

Defensible imagery scale

23.22 trillion four-channel image values.

The number is large because each of 115.71 million road-centered chips was standardized to a 224 × 224 tensor and evaluated across red, green, blue, and near-infrared. It is not a count of unique photographs or unique source pixels.

115.71 millionroad-centered image chips
×
224 × 224tensor locations per chip
×
4 channelsRGB + near-infrared
=
23.22 trillion
One tensor value is one red, green, blue, or near-infrared number at one standardized image location. Multiplying four channels across 224 by 224 locations and 115.71 million chips produces 23.22 trillion processed values.

Exact calculation: 115,711,736 × 224 × 224 × 4 = 23,223,808,262,144. Separate red/NIR-derived spectral measurements are additional work and are not added to this headline.

The vision branch

65.02 millionroad sections represented in the final EfficientNet artifacts

Road-centered RGB and near-infrared chips feed spatially separated supervised models, while vision-language and self-supervised representations add different visual perspectives. Their outputs become road-level inputs to the larger system.

Structured fusion beyond vision

300.54 billiontotal feature-value slots across the complete five-region production matrices

Within those matrices, OSM road attributes and topology, Overture buildings and land context, road-name embeddings, terrain, water, vegetation, agriculture, soils, climate, development, nighttime light, and routing-derived road importance contribute alongside imagery.

15.05B routing attempts5.36B contributing paths8,401 active feature columns
02

Scenic intelligence

How Scenic Is the View from the Road?

Rank routes and road networks by modeled visible experience, not by nearby attractions alone. Sherpa tests what an eye-height traveler can actually see after terrain, vegetation, buildings, and other occluders are considered.

Explore scenic data

Scenic intelligence

Model what is visible from the road, not what happens to be nearby.

Sherpa samples the road at eye height and looks outward through terrain, vegetation, water, structures, and land context. A lake behind a warehouse, a monument behind trees, or a mountain behind a ridge does not count as though it were unobstructed.

Scenic road-ranking layer over processed Sentinel-2 imagery
Lower modeled scenic contextHigher modeled scenic context
A road-level comparative experience layer, built for ranking, discovery, tourism, outdoor products, and route selection. Contains modified Copernicus Sentinel data (2019–2025). Road data © OpenStreetMap contributors; Overture Maps.
64.31Mmapped road sections in the scenic layer
1.624Broad sample locations
97.456B
11purpose-built visibility/context rasters

The result is a consistent road-to-road scenic ranking for discovery, premium route choice, touring, outdoor navigation, tourism, and destination products across an entire network.

A configured sightline is one eye-height ray cast from a sampled road location in a selected horizontal direction and pitch band. It tests the first terrain, water, vegetation, structure, or land-context feature that blocks or defines the view.
Five aligned Great River Road panels showing Sentinel imagery, terrain, water, vegetation and buildings, and representative scenic sightlines
One Great River Road corridor, reconstructed five ways: the panels share the same Wisconsin footprint and follow the same road through recent multispectral imagery, 188 to 349 meter terrain, Mississippi River water evidence, vegetation and real Overture building footprints, then a legible subset of the 14,100 stored first-hit sightlines. The completed scene includes 235 eye-height road samples, 3,058 water first-hits, 7,995 terrain first-hits, and building occlusion checks. Each layer answers a different part of the same practical question: what can a traveler actually see?

How Sherpa measures scenic value

Every 25 meters, we test what a traveler can physically see from the road.

Sherpa places a virtual observer 1.7 meters above the road and casts 60 lines of sight through a 3D model of the surrounding landscape. Each line identifies the first visible feature it reaches. Mountain relief, water, forests, open land, landmarks, and historic structures can contribute scenic value. A nearer ridge, tree, building, warehouse, or other obstruction can block what lies behind it.

The practical difference: a lake behind a warehouse is not treated like an open lake view, and a mountain hidden behind a nearer ridge is not counted as visible. The score reflects the experience from the road, not a list of attractions somewhere nearby.

25 m

Move along the road

Road geometry is sampled every 25 meters.

1.7 m

Use normal eye height

The virtual observer stands 1.7 meters above modeled road elevation.

20 × 3

Cast 60 lines of sight

Twenty directions across 220° are tested at three vertical viewing angles.

30 km

Follow near and distant views

Terrain-aware sightlines can extend to 30 kilometers where the landscape permits.

Pikes Peak Highway, Colorado

The green lines show representative views tested from the road.

The black path is Pikes Peak Highway. Each green line begins at an eye-height road sample and travels through the 3D landscape until it reaches the first visible terrain, water, vegetation, structure, or land-context feature.

Pikes Peak terrain scene showing the highway and green line-of-sight checks cast from eye-height road samples
What this example shows: Sherpa checks the view a driver, rider, or runner could physically experience while climbing Pikes Peak Highway. Only a representative subset of lines is drawn here so the scene remains readable. The production scenic layer applies the same visibility logic across 1.624 billion road sample locations and 97.456 billion configured sightlines.

Terrain and landform

Relief, prominence, horizon, ridges, cliffs, peaks, valleys, rock, beaches, dunes, and open landscape context.

Water and shoreline

Sentinel-2 RGB/red/NIR-derived masks, terrain plausibility, and vector-water continuity support lakes, rivers, reservoirs, wetlands, and coastal views.

Vegetation and land cover

Low and tall vegetation, available height context, forests, parks, reserves, meadows, wetlands, orchards, and vineyards.

Structures and built context

Overture buildings and available height/floor context act as occluders. Upstream civic, historic, religious, tower, bridge, museum, monument, and castle-like classes can contribute where identified.

Negative built context

Industrial, retail, construction, airport, landfill, quarry, port, military, and other built clutter can reduce the experience or block what lies behind it.

Road aggregation and quality control

Point-level observations become road scores with support counts and diagnostics. Regional checks verify raster alignment, score continuity, support density, and finished delivery ranges.

See the finished score and the method.Use the live atlas for continental road values, then open the 3D theater to inspect the visibility logic.
03

Modeled traffic

How Heavily Is the Road Likely to Be Used?

Add a comparative road-use estimate where live counts, vehicle telemetry, or proprietary athlete activity are unavailable. Rank roads within the network for likely use, importance, and avoidance decisions.

Explore modeled traffic

Modeled traffic

From nighttime activity to modeled traffic on 66.22 million road sections.

Recent nighttime light identifies likely centers of human activity. Sherpa routes billions of plausible local, commuter, and intercity journeys between them on the real road network, then measures which roads appear repeatedly. That repeated use becomes modeled traffic.

The three views below show one continuous process across the same regional footprint. First, locate likely journey origins and destinations. Next, route many plausible journeys between them. Finally, accumulate repeated road use into the modeled traffic layer.

01

Find likely activity centers

Where journeys are likely to begin and end.

Recent VIIRS nighttime light and retained likely journey origins and destinations over the road network

Recent VIIRS nighttime light reveals likely centers of human activity and weights where modeled journeys begin and end.

02

Route plausible journeys

Many routes connect those centers.

Many representative RoutingKit routes connecting exact retained activity points across the road network

The blue lines are 48 real representative RoutingKit paths between exact retained activity points, spanning local, regional, and longer journeys.

03

Build modeled traffic

Repeated road use becomes a comparable score.

Finished modeled traffic layer showing accumulated relative road use across the same road network
Lower likely road useHigher likely road use

Roads rank higher when they repeatedly carry more modeled journeys. The result supports routing, avoidance, network analysis, and map enrichment.

15.05Borigin-destination attempts
5.36Busable routed paths
66.22Mroad sections with modeled traffic

From activity to a road-level product: production runs applied this process at continental scale. Each usable path contributed to the roads it traveled, and billions of those contributions became the finished modeled traffic ranking.

What this means: modeled traffic estimates relative road use where direct measurements are unavailable. It is not live congestion, observed vehicle counts, or customer trace data.

Modeled traffic values over processed Sentinel-2 imagery
Lower likely road useHigher likely road use
The map shows modeled traffic across the same Colorado footprint as the other layer examples. Contains modified Copernicus Sentinel data (2019–2025). Road data © OpenStreetMap contributors; Overture Maps.

More than shortest-path counting.

Nighttime-light cells identify likely centers of human activity. The sampler mixes local, commute-like, and intercity journeys, with land-use context available for residential-to-commercial patterns. RoutingKit make continental query volume practical.

Successful paths are trimmed near their endpoints, accumulated with road-structure weighting, and repeated in masking rounds so dominant highways do not hide meaningful secondary roads. Tie-aware percentile normalization creates a comparable road rank.

Inspect modeled traffic in the live atlas
A contraction hierarchy is a preprocessed road graph that answers shortest-path queries much faster while preserving exact route distance under the configured weights. That speed made billions of continental routing attempts practical.
Activity proxyHarmonized 2024 nighttime-light cells

Brightness-aware candidates and spatial reduction identify likely journey centers.

Journey mixLocal, commute-like, and intercity samples

Pairs are sampled. They are not claimed as every unique origin/destination combination.

Fast routingC++ and RoutingKit contraction hierarchies

Continental graphs support billions of shortest-path attempts.

Road accumulationEndpoint trimming, weighting, and masking rounds

Usable path contributions become a road-level comparative rank.

Independent activity context

Recent nighttime light, not customer traces.

Harmonized 2024 VIIRS brightness locates likely centers of activity. Sherpa turns those centers into a controlled mix of local, commute-like, and intercity origin-destination samples without requiring private athlete or vehicle telemetry.

Purpose-built throughput

A routing system built for billions of queries.

The released C++ simulator uses RoutingKit contraction hierarchies, endpoint trimming, structural weighting, masking rounds, and tie-aware normalization to convert billions of path contributions into a stable road-level rank.

One geospatial intelligence platform

The datasets are built to act together.

Surface, scenery, modeled traffic, elevation, geometry, and route context are attached to one high-performance graph. The same attributes can generate a route, explain it, match an imported route, or enrich a partner's own network.

Road intelligenceSurface classificationScenic rankingModeled trafficElevation and road context
Road IDs and per-arc attributes
Shared graph92.71M nodes240.13M directed arcs1.275B geometry pointsFresh production bike-graph snapshot; values may change with releases.
Search, match, score, refine
Product surfacesRoute generationRoute IntelligenceRoute Studio + Sherpa MapAPI and MCP alpha

Experience-aware route generation

More than point-to-point shortest path.

The engine can target distance, surface composition, scenery, elevation, steepness, overlap, backtracking, turns, curvature, traffic exposure, road/path preference, required roads, destinations, and route ingredients.

Different request modes use simulated annealing, genetic mutation and crossover, weighted pathfinding, and refinement. They are selectable search paths, not a fixed sequence forced onto every request.

Route Intelligence

Generate, import, inspect, and revise.

Route Intelligence shares the loaded graph to search roads, match uploaded GPX, TCX, FIT, GeoJSON, and route files, find water, parking, and POIs, materialize required route ingredients, and aggregate surface, scenic, elevation, hill, and flow context.

The data therefore does two jobs: it powers route creation and remains available to explain what the route contains.

Production

Surface, scenic, modeled traffic

Finished continental road-level datasets with live visualization and delivery fields.

Production support

Elevation atlas and route graph

102.56 million atlas ways, 1.148 billion elevation points, and shared per-arc intelligence.

Operational indexes

POI, water, structures, and ingredients

Graph-linked support for route generation, matching, analysis, and product experiences.

Active development

Climb intelligence and MCP/API

Substantial systems under active release work; not presented as finished commercial datasets.

Operating proof

Already powering live Sherpa products.

The data is not waiting for a demo application. Route Studio and Sherpa Map use Sherpa road intelligence in real routing and planning experiences. Partners can license the data without adopting either interface.

Prompt-to-route alpha

Natural language can now drive the same route intelligence.

The WIP public alpha coordinates separate Route Generation and Route Analysis MCP services. It can create a route from distance, region, start/end, and experience preferences. It can also match an uploaded route and keep revising it in the same workspace.

  • Create routes around target distance and rider, runner, or driver preferences.
  • Ask questions, generate maps and graphs, and run deeper route analyses.
  • Revise a generated or uploaded route through the same conversation.
Try the WIP demoOpen the API/MCP portal

Public alpha; availability may vary during active development. Video runs 4:56, contains no audio, and speeds up selected processing sequences. The demo is proof of integration capability, not a required data-delivery interface.

Commercial use and delivery

Fit the data to the business, not the business to a fixed endpoint.

Start with a graph, geography, or product decision. Sherpa can return partner-keyed attributes, geometry, tiles, a hosted layer, or another agreed structure with clear field definitions and provenance.

01

Graph-matched attributes

Return Sherpa fields against partner edge IDs, supplied geometry, or an agreed conflation key.

02

Common geospatial formats

Parquet, GeoParquet, CSV, FlatGeobuf, GIS packages, vector tiles, hosted layers, or a custom schema.

03

Confidence and support policy

Set high/medium/low bands, unknown handling, review thresholds, and coverage-versus-certainty tradeoffs for the use case.

04

Provenance-ready package

Use a license-cleared source/model profile with release notes, field definitions, source attribution, and a delivery manifest.

Consumer and navigation

Routing, fitness, automotive, outdoor, and tourism

Surface-aware selection, warnings, discovery, scenic alternatives, exposure avoidance, route summaries, and premium experiences.

Maps and data platforms

Graph completion, QA, and differentiated attributes

Standardize road context, prioritize review, enrich basemaps, compare networks, and add proprietary road-level layers.

Enterprise and GIS

Planning, prioritization, and spatial analysis

Corridor comparison, access analysis, tourism planning, network ranking, infrastructure review, and scenario support.

A practical first engagement

One road set. One baseline. One measurable decision.

Choose a representative geography and a product question worth improving. Sherpa will fit a sample to the identifiers, fields, confidence policy, and format needed for a direct comparison.

  1. 01Select the decision and accepted baseline.
  2. 02Match the graph, geography, and fields.
  3. 03Receive a purpose-built evaluation package.
  4. 04Measure coverage, agreement, ranking, utility, or product lift.

Direct contact

Put the data against your current road set.

Bring one valuable decision and one representative geography. We will define a clear evaluation that your technical team can verify.

Explore live data

Curated evaluation data

Request a sample built around your decision.

Describe one geography and one product question. Sherpa will respond with a practical sample scope, useful fields, and a delivery format your team can evaluate directly.

Your submission goes directly to Sherpa's private admin inbox. Eric Semianczuk will review it and reply to the work email above. Do not paste proprietary road files here; if files are needed, Eric will provide a secure next step.