Commercial road-surface intelligence · North America + Europe
96.09% of roads across North America and Europe classified as paved or unpaved.
License confidence-scored paved or unpaved road-surface classifications. Sherpa classifies all 66.40 million production road sections using multispectral imagery and broad geographic context. With a commitment to flexibility, data delivery is matched to your requirements.
Road Surface Classification is the lead product in a broader road-data portfolio, alongside Scenic Road Score and Modeled Traffic, a consistent comparative network-use metric.
Validated performance: 94.4% and 86.9% accuracy across 1.95 million imagery predictions.
ROC-AUC measures how reliably the imagery model ranks paved examples above unpaved examples across every possible decision threshold. A score of 50% would be random ranking; Sherpa achieved 94.4%. Out-of-fold means each benchmark prediction came from a model that was not trained on that road's local geographic cell. This makes the test harder to pass through geographic memorization.
Three finished road datasets
Road Surface, Scenic Value, and Modeled Traffic. Defined.
Three direct product questions: Is the road paved? What can a traveler see from it? How heavily is it likely to be used? Each finished answer is backed by continental imagery, graph engineering, terrain analysis, geographic context, machine learning, and production delivery infrastructure.
Production dataset01 · Lead product
Road Surface: Paved or Unpaved
All 66.40 million production road sections are independently classified with confidence, including 41.32 million that had no usable OpenStreetMap paved or unpaved tag.
Production dataset02 · Scenic intelligence
How Scenic Is the View from the Road?
A terrain- and occlusion-aware ranking built from eye-height views of water, landforms, vegetation, structures, and built context.
Production dataset03 · Modeled traffic
How Heavily Is the Road Likely to Be Used?
A comparative estimate of road use built from billions of sampled routing attempts. It is not live traffic, observed counts, or athlete activity.
Lead data product
Road Surface: Paved or Unpaved
Give routing, mapping, fitness, mobility, and outdoor products one consistent surface answer across the full production network. Sherpa classifies roads with missing tags and reclassifies tagged roads using newer imagery and geographic context.
The commercial gap
OSM supplied usable surface tags for 37.8%. Sherpa independently classifies the full network.
Across the same 66.40 million North American and European production road sections, 25.08 million carried a usable OSM paved or unpaved source tag. Sherpa does not preserve those tags as final answers. It reclassifies all 66.40 million sections using newer imagery, road and network context, language, land use, structures, terrain, development, and modeled road importance. Of those roads, 41.32 million, or 62.2%, had no usable OSM surface tag at all.
Same production network, same denominator
OSM source coverage compared with Sherpa classification coverage
This measures whether a road had usable source metadata before Sherpa. OSM tags support training and evaluation where present, but the finished Sherpa layer reclassifies those roads too. It does not copy 37.8% and fill only the remainder.
The engineering beneath one road color
The visible answer is only the tip.
One paved or unpaved value can draw on multispectral imagery, language, road geometry, graph structure, terrain, land context, buildings, and billions of modeled journeys. No single input decides the answer. The depth comes from independent measurements that reinforce or correct one another.
Structured fusion
Specialized model stages
Engineered measurements
Continental source systems
Technical detailHow the imagery, model stages, text, and structured fusion work
Imagery preparation
Sentinel-2 Global Mosaic RGB, B04 red, and B08 near-infrared are sampled around roads. B04 is aligned to the NIR reference grid where needed. The chipper applies a road mask and nodata/validity filters, then measures center-road and along-road red, NIR, NDVI, DVI, simple ratio, brightness, RGB ratios, saturation, excess green, quantiles, threshold shares, and gradients.
Vision sequence
Five geographically grouped folds produce untouched stage-one predictions. Production model stages aggregate per-chip probabilities, uncertainty measurements, and learned representations into road-level evidence. The finished feature stacks combine convolutional predictions with vision-language and self-supervised representations where available; these become inputs to the structured model rather than a vote-only ensemble.
Road-name encoding
Primary, alternate, official, local, English, address-street, and reference fields are normalized and deduplicated. A pretrained sentence encoder creates 768-dimensional representations, reduced and cached as 64 numerical dimensions so names can inform the model without millions of sparse categories.
Final fusion
CatBoost combines numerical, categorical, image, text, graph, and geographic representations into the finished road-level decision. Each delivered classification includes the Sherpa class, confidence score, and whether a usable OSM source tag existed before reclassification. Tagged roads receive spatially out-of-fold predictions; untagged roads receive full-model inference. Counts use one mapped road section per distinct OSM way identifier, not one unique named road or routing graph edge.
Completed surface benchmark
Proven road surface classification.
Evidence: 1,947,588 balanced road predictions across five . Roads in the same H3 resolution-6 cell remain in the same fold, so the same local cell cannot appear in both training and validation.
Scope: this benchmark measures the spatially held-out imagery stage on its own. The finished production classification adds the structured geographic, graph, language, and routing context shown above.
Confidence-ranked coverage
Higher-confidence imagery decisions are substantially more accurate.
Selective results from the same untouched benchmark. is distance from the 0.5 decision boundary, not a claim of calibrated probability.
Defensible imagery scale
23.22 trillion four-channel image values.
The number is large because each of 115.71 million road-centered chips was standardized to a 224 × 224 tensor and evaluated across red, green, blue, and near-infrared. It is not a count of unique photographs or unique source pixels.
Exact calculation: 115,711,736 × 224 × 224 × 4 = 23,223,808,262,144. Separate red/NIR-derived spectral measurements are additional work and are not added to this headline.
The vision branch
65.02 millionroad sections represented in the final EfficientNet artifactsRoad-centered RGB and near-infrared chips feed spatially separated supervised models, while vision-language and self-supervised representations add different visual perspectives. Their outputs become road-level inputs to the larger system.
Structured fusion beyond vision
300.54 billiontotal feature-value slots across the complete five-region production matricesWithin those matrices, OSM road attributes and topology, Overture buildings and land context, road-name embeddings, terrain, water, vegetation, agriculture, soils, climate, development, nighttime light, and routing-derived road importance contribute alongside imagery.
Scenic intelligence
How Scenic Is the View from the Road?
Rank routes and road networks by modeled visible experience, not by nearby attractions alone. Sherpa tests what an eye-height traveler can actually see after terrain, vegetation, buildings, and other occluders are considered.
Scenic intelligence
Model what is visible from the road, not what happens to be nearby.
Sherpa samples the road at eye height and looks outward through terrain, vegetation, water, structures, and land context. A lake behind a warehouse, a monument behind trees, or a mountain behind a ridge does not count as though it were unobstructed.
The result is a consistent road-to-road scenic ranking for discovery, premium route choice, touring, outdoor navigation, tourism, and destination products across an entire network.
How Sherpa measures scenic value
Every 25 meters, we test what a traveler can physically see from the road.
Sherpa places a virtual observer 1.7 meters above the road and casts 60 lines of sight through a 3D model of the surrounding landscape. Each line identifies the first visible feature it reaches. Mountain relief, water, forests, open land, landmarks, and historic structures can contribute scenic value. A nearer ridge, tree, building, warehouse, or other obstruction can block what lies behind it.
The practical difference: a lake behind a warehouse is not treated like an open lake view, and a mountain hidden behind a nearer ridge is not counted as visible. The score reflects the experience from the road, not a list of attractions somewhere nearby.
Move along the road
Road geometry is sampled every 25 meters.
Use normal eye height
The virtual observer stands 1.7 meters above modeled road elevation.
Cast 60 lines of sight
Twenty directions across 220° are tested at three vertical viewing angles.
Follow near and distant views
Terrain-aware sightlines can extend to 30 kilometers where the landscape permits.
Pikes Peak Highway, Colorado
The green lines show representative views tested from the road.
The black path is Pikes Peak Highway. Each green line begins at an eye-height road sample and travels through the 3D landscape until it reaches the first visible terrain, water, vegetation, structure, or land-context feature.
Terrain and landform
Relief, prominence, horizon, ridges, cliffs, peaks, valleys, rock, beaches, dunes, and open landscape context.
Water and shoreline
Sentinel-2 RGB/red/NIR-derived masks, terrain plausibility, and vector-water continuity support lakes, rivers, reservoirs, wetlands, and coastal views.
Vegetation and land cover
Low and tall vegetation, available height context, forests, parks, reserves, meadows, wetlands, orchards, and vineyards.
Structures and built context
Overture buildings and available height/floor context act as occluders. Upstream civic, historic, religious, tower, bridge, museum, monument, and castle-like classes can contribute where identified.
Negative built context
Industrial, retail, construction, airport, landfill, quarry, port, military, and other built clutter can reduce the experience or block what lies behind it.
Road aggregation and quality control
Point-level observations become road scores with support counts and diagnostics. Regional checks verify raster alignment, score continuity, support density, and finished delivery ranges.
Modeled traffic
How Heavily Is the Road Likely to Be Used?
Add a comparative road-use estimate where live counts, vehicle telemetry, or proprietary athlete activity are unavailable. Rank roads within the network for likely use, importance, and avoidance decisions.
Modeled traffic
From nighttime activity to modeled traffic on 66.22 million road sections.
Recent nighttime light identifies likely centers of human activity. Sherpa routes billions of plausible local, commuter, and intercity journeys between them on the real road network, then measures which roads appear repeatedly. That repeated use becomes modeled traffic.
The three views below show one continuous process across the same regional footprint. First, locate likely journey origins and destinations. Next, route many plausible journeys between them. Finally, accumulate repeated road use into the modeled traffic layer.
Find likely activity centers
Where journeys are likely to begin and end.
Recent VIIRS nighttime light reveals likely centers of human activity and weights where modeled journeys begin and end.
Route plausible journeys
Many routes connect those centers.
The blue lines are 48 real representative RoutingKit paths between exact retained activity points, spanning local, regional, and longer journeys.
Build modeled traffic
Repeated road use becomes a comparable score.
Roads rank higher when they repeatedly carry more modeled journeys. The result supports routing, avoidance, network analysis, and map enrichment.
From activity to a road-level product: production runs applied this process at continental scale. Each usable path contributed to the roads it traveled, and billions of those contributions became the finished modeled traffic ranking.
What this means: modeled traffic estimates relative road use where direct measurements are unavailable. It is not live congestion, observed vehicle counts, or customer trace data.
More than shortest-path counting.
Nighttime-light cells identify likely centers of human activity. The sampler mixes local, commute-like, and intercity journeys, with land-use context available for residential-to-commercial patterns. RoutingKit make continental query volume practical.
Successful paths are trimmed near their endpoints, accumulated with road-structure weighting, and repeated in masking rounds so dominant highways do not hide meaningful secondary roads. Tie-aware percentile normalization creates a comparable road rank.
Inspect modeled traffic in the live atlasBrightness-aware candidates and spatial reduction identify likely journey centers.
Pairs are sampled. They are not claimed as every unique origin/destination combination.
Continental graphs support billions of shortest-path attempts.
Usable path contributions become a road-level comparative rank.
Independent activity context
Recent nighttime light, not customer traces.
Harmonized 2024 VIIRS brightness locates likely centers of activity. Sherpa turns those centers into a controlled mix of local, commute-like, and intercity origin-destination samples without requiring private athlete or vehicle telemetry.
Purpose-built throughput
A routing system built for billions of queries.
The released C++ simulator uses RoutingKit contraction hierarchies, endpoint trimming, structural weighting, masking rounds, and tie-aware normalization to convert billions of path contributions into a stable road-level rank.
One geospatial intelligence platform
The datasets are built to act together.
Surface, scenery, modeled traffic, elevation, geometry, and route context are attached to one high-performance graph. The same attributes can generate a route, explain it, match an imported route, or enrich a partner's own network.
Experience-aware route generation
More than point-to-point shortest path.
The engine can target distance, surface composition, scenery, elevation, steepness, overlap, backtracking, turns, curvature, traffic exposure, road/path preference, required roads, destinations, and route ingredients.
Different request modes use simulated annealing, genetic mutation and crossover, weighted pathfinding, and refinement. They are selectable search paths, not a fixed sequence forced onto every request.
Route Intelligence
Generate, import, inspect, and revise.
Route Intelligence shares the loaded graph to search roads, match uploaded GPX, TCX, FIT, GeoJSON, and route files, find water, parking, and POIs, materialize required route ingredients, and aggregate surface, scenic, elevation, hill, and flow context.
The data therefore does two jobs: it powers route creation and remains available to explain what the route contains.
Surface, scenic, modeled traffic
Finished continental road-level datasets with live visualization and delivery fields.
Elevation atlas and route graph
102.56 million atlas ways, 1.148 billion elevation points, and shared per-arc intelligence.
POI, water, structures, and ingredients
Graph-linked support for route generation, matching, analysis, and product experiences.
Climb intelligence and MCP/API
Substantial systems under active release work; not presented as finished commercial datasets.
Operating proof
Already powering live Sherpa products.
The data is not waiting for a demo application. Route Studio and Sherpa Map use Sherpa road intelligence in real routing and planning experiences. Partners can license the data without adopting either interface.

Route Studio
Uses proprietary surface and scenic intelligence to shape route choice, then carries those decisions into route analysis and navigation.
Visit Route Studio

Sherpa Map
Exposes surface-aware route preferences and road context directly in a public web route builder.
Open Sherpa MapPrompt-to-route alpha
Natural language can now drive the same route intelligence.
The WIP public alpha coordinates separate Route Generation and Route Analysis MCP services. It can create a route from distance, region, start/end, and experience preferences. It can also match an uploaded route and keep revising it in the same workspace.
- Create routes around target distance and rider, runner, or driver preferences.
- Ask questions, generate maps and graphs, and run deeper route analyses.
- Revise a generated or uploaded route through the same conversation.
Public alpha; availability may vary during active development. Video runs 4:56, contains no audio, and speeds up selected processing sequences. The demo is proof of integration capability, not a required data-delivery interface.
Commercial use and delivery
Fit the data to the business, not the business to a fixed endpoint.
Start with a graph, geography, or product decision. Sherpa can return partner-keyed attributes, geometry, tiles, a hosted layer, or another agreed structure with clear field definitions and provenance.
Graph-matched attributes
Return Sherpa fields against partner edge IDs, supplied geometry, or an agreed conflation key.
Common geospatial formats
Parquet, GeoParquet, CSV, FlatGeobuf, GIS packages, vector tiles, hosted layers, or a custom schema.
Confidence and support policy
Set high/medium/low bands, unknown handling, review thresholds, and coverage-versus-certainty tradeoffs for the use case.
Provenance-ready package
Use a license-cleared source/model profile with release notes, field definitions, source attribution, and a delivery manifest.
Consumer and navigation
Routing, fitness, automotive, outdoor, and tourism
Surface-aware selection, warnings, discovery, scenic alternatives, exposure avoidance, route summaries, and premium experiences.
Maps and data platforms
Graph completion, QA, and differentiated attributes
Standardize road context, prioritize review, enrich basemaps, compare networks, and add proprietary road-level layers.
Enterprise and GIS
Planning, prioritization, and spatial analysis
Corridor comparison, access analysis, tourism planning, network ranking, infrastructure review, and scenario support.
A practical first engagement
One road set. One baseline. One measurable decision.
Choose a representative geography and a product question worth improving. Sherpa will fit a sample to the identifiers, fields, confidence policy, and format needed for a direct comparison.
- 01Select the decision and accepted baseline.
- 02Match the graph, geography, and fields.
- 03Receive a purpose-built evaluation package.
- 04Measure coverage, agreement, ranking, utility, or product lift.
Direct contact
Put the data against your current road set.
Bring one valuable decision and one representative geography. We will define a clear evaluation that your technical team can verify.