Overview
Airborne LiDAR surveys routinely produce single files of several gigabytes and tens of millions of points, a volume that cannot be transferred to a web browser as a unit and cannot be rendered by conventional graphics pipelines without discarding the vast majority of the data at any given moment.
The system described here accepts a LAS file as input and produces a tiled representation suitable for range-limited HTTP delivery from a CDN, a browser client that loads a coarse overview immediately and refines it without blocking, and an automated ingestion path that converts a dropped file into a published, browsable dataset without further operator action.
"Streaming," in the sense used throughout this report, does not mean a continuous byte transfer of the whole file. It means restructuring the dataset once, ahead of time, into a hierarchy of spatial tiles ordered from coarse to fine, so that a client can request only the tiles relevant to its current view and refine that view as the camera approaches.
Background
LAS and LAZ
LAS is a binary file format for storing point-cloud data, standardized by ASPRS. Each point occupies a fixed-size record defined by a "point data record format" (PDRF) number. LAZ is a lossless compressed variant. Neither format is tiled or organized by level of detail: reading any region requires scanning some or all of the file.
The Octree and Level-of-Detail
An octree is a tree structure in which each node represents a cubical region of 3D space, and each non-leaf node's cube is subdivided into eight equal child cubes. Applied to a point cloud, an octree node holds a spatially thinned subset of the points that fall within its cube. The root encloses the entire survey; its children hold finer subsets of smaller sub-volumes.
This structure enables a level-of-detail (LOD) scheme: a renderer displays the root node alone as a coarse representation, then descends into whichever child nodes fall within the camera's view and are worth resolving, replacing coarseness with detail only where a viewer would notice it.
The Entwine Point Tile Format
EPT is a format and directory layout for storing a point cloud as an octree of binary tiles, produced by the Entwine tool. An EPT dataset consists of a manifest file (ept.json) describing bounds, schema, and structure, together with one binary data file per node and a set of hierarchy files describing which nodes exist.
The Role of a CDN
A content delivery network caches static files at edge locations distributed relative to end users, so repeated requests for the same file are served from a nearby cache rather than from the origin on every request. For a tiled point cloud, where coarse tiles are requested by every viewer and fine tiles are requested repeatedly by users zooming into the same regions, this cache behaviour keeps load times and origin bandwidth costs bounded as viewership grows.
Source Data Characteristics
Before any conversion step, the source LAS file should be inspected to confirm the properties that later pipeline stages assume. Inspection was performed with PDAL, run directly against the source file:
pdal info --summary source.las
The properties that matter for the stages that follow are the point count and file size, the point data record format, the coordinate reference system, and the spatial extent.
| Property | Value |
|---|---|
| Point Count | 91,625,792 |
| File Size | 2.9 GB |
| PDRF | 7 (XYZ, RGB, GPS time) |
| Ground Footprint | 367 m × 271 m |
| Vertical Extent | ~30 m |
Environment
Three tools are required before any conversion: Entwine to build the EPT octree, PDAL to inspect and decimate the source file, and a Python 3.13 environment with boto3 and laspy for the publication stage. Entwine was installed in an isolated conda environment, because it links its own build of GDAL and PDAL, and installing it alongside other GIS software risks conflicting shared libraries.
conda create -n atlas-pointcloud -c conda-forge entwine conda activate atlas-pointcloud pip install boto3 laspy
ModuleNotFoundError: No module named '_gdal'. The reported error gives no indication that the actual cause is PATH ordering. Activate the environment and confirm with where entwine (Windows) or which entwine (Unix).Conversion to EPT
With the environment active, the EPT octree is built from the source LAS file:
entwine build -i source.las -o output/ept --dataType binary
The --dataType binary flag is not optional for this pipeline. The client renderer decodes raw binary node records directly; it does not implement a LAZ decompressor, and a node written in Entwine's laszip data type is rejected outright. Passing the flag explicitly means a future change to Entwine's default output format cannot silently produce a dataset the existing renderer is unable to read.
For the representative dataset, this build produced 1,914 octree nodes. Each node holds points sampled on a grid of span voxels per edge of its cube, where span is 128.
Output Directory
Manifest: bounds, schema, point count
One binary file per octree node
JSON files describing node existence and tree connections
ept-hierarchy/<node-id>.json, which the client must fetch in turn. A client that reads only the root file will build a tree that terminates at whatever depth the root describes directly. This produces no error: the symptom is a dataset that renders only its coarsest levels, indistinguishable from one genuinely built to a shallow depth.Payload Optimisation
Entwine's default schema carries 22 dimensions at 49 bytes per point. The renderer consumes only six dimensions (X, Y, Z, Red, Green, Blue), meaning 62 percent of every transferred byte is discarded on arrival. The schema was restricted to these six dimensions:
entwine build \ -i source.las \ -o output/ept \ --dataType binary \ --noOriginId \ --config schema-override.json
{
"schema": [
{ "name": "X" },
{ "name": "Y" },
{ "name": "Z" },
{ "name": "Red" },
{ "name": "Green" },
{ "name": "Blue" }
]
}Restricting the schema reduced the per-point record from 49 to 18 bytes, a factor of 2.72. Decoded geometry and colour from the restricted build were verified byte-identical to the corresponding values from the unrestricted build.
| Metric | Default Schema | Restricted Schema | Reduction |
|---|---|---|---|
| Bytes per point | 49 B | 18 B | 2.72x |
| Total node data | 4,194 MB | ~1,541 MB | 2.72x |
| Median node file | ~2 MB | ~760 kB | 2.63x |
Publication
The built node tree is synchronized to object storage under a per-project prefix and served through a CDN:
aws s3 sync output/ept \ s3://<S3_BUCKET_NAME>/projects/<PROJECT_SLUG>/pointcloud/ept/
Publication emits a machine-readable summary containing the public URL of the published ept.json, which is written back to the corresponding project's database record so the client application knows where to fetch the dataset without further manual configuration.
Automation
A manual build-and-publish procedure does not scale past the first survey. The pipeline is triggered by a local worker process that watches an input directory; dropping a file named pointcloud.las into a project's folder enqueues a build job.
Completed work is recorded as a fingerprint of the file's size and modification time at the moment it was processed, so restarting the worker does not reprocess every file already present. Fingerprints are recorded for failed builds as well as successful ones; without this, a file that reliably crashes the builder would be retried indefinitely on every restart rather than surfaced once as a failure requiring operator attention.
Client Architecture
The browser-side renderer is built on three.js, using a single THREE.Points object per loaded octree node and one shared PointsMaterial across all of them. A custom scheduler fetches the manifest, walks the full hierarchy including continuation files, renders the root node immediately so a coarse view appears without waiting for deeper fetches, then progressively requests finer nodes as the camera and the loading gate permit.
World-space coordinates require float32 precision
256 levels per channel is sufficient; GPU normalizes to 0-1
This reduces the colour attribute from 12 bytes per point (three 32-bit floats) to 3 bytes, and the combined per-point GPU footprint from 24 to 15 bytes, a reduction of 37.5 percent, while the vertex shader still receives colour values in the 0-to-1 range it expects.
Decoding EPT Coordinates
EPT stores each point's coordinates not as raw floating-point positions but as scaled integers: the true coordinate is raw × scale + offset, where scale and offset are given per dimension in the schema.
bounds field in ept.json describes the octree's enclosing cube, which is necessarily larger than the surveyed data. A separate field, boundsConforming, gives the true, non-cubical extent of the data. Centring the scene on the enclosing cube rather than on boundsConforming placed roughly half of the point cloud below the intended reference grid.Performance Engineering
The governing principle: a point the screen cannot resolve should never be retained in memory or on the GPU. Each node samples its cube on a grid with span divisions per edge, so adjacent points sit edge / span apart in world units. Projecting this spacing to screen pixels, given the node's distance from the camera, tests directly whether loading that node adds visible detail.
Three Corrected Defects
- 01
No threshold on loading
The scheduler queued every node in the frustum with no threshold on whether loading it would produce a visible change. Each camera movement pulled in strictly more data than the last, until the entire dataset was resident. The symptom was deceptive: performance was acceptable for the first few seconds and degraded progressively thereafter.
- 02
Eviction considered only out-of-frustum nodes
Because nearly every loaded node remains inside the frustum at any given moment, this left no eviction candidates, meaning the point budget could not be enforced at all. It also rendered the adaptive frame-rate tuner inert.
- 03
LRU timestamp never updated
Each node's lastAccess timestamp was written once, at insertion, and never updated on subsequent frames. LRU ordering was equivalent to plain insertion order, so eviction removed nodes unrelated to how recently they had been rendered.
Verification Methodology
Two standalone Node.js scripts assert directly against the live, published dataset rather than against local fixtures, since the defects described above are properties of the full-scale, deployed data and its network delivery.
Asserts on the spread and bounds of decoded point coordinates, not merely the type or shape of the buffer
Drives 60 simulated camera movements and asserts the resident count converges rather than climbing without bound
Results
The pipeline transforms a 2.9 GB LAS file into a CDN-served EPT dataset that a browser can render interactively. The schema restriction cuts transferred bytes by 2.72x, the client keeps the resident point count between one and two orders of magnitude below the source dataset, and the pixel-resolution loading gate ensures a point the screen cannot resolve is never retained.
| Metric | Before | After |
|---|---|---|
| Per-point payload | 49 B | 18 B |
| Total node data | 4,194 MB | ~1,541 MB |
| Resident points (wide view) | Unbounded | 3.07 M pts |
| GPU per-point footprint | 24 B | 15 B |
Limitations and Future Work
No compression
The --dataType binary requirement forgoes Entwine's laszip output. A compressed format compatible with client-side decoding is a possible further avenue.
Single-worker automation
The automation is a single local worker watching one input path per project. Its behaviour under simultaneous uploads across projects, or recovery after a crash mid-build, is not yet covered.
Platform-specific PATH issues
The Windows-specific PATH pitfalls were observed on that platform; equivalent pitfalls on macOS or Linux are not addressed.
Fixed device-pixel-ratio cap
The device-pixel-ratio cap is a fixed 1.5 chosen for the hardware on which it was measured. An adaptive cap responding to a monitored frame budget is a plausible extension.
Reproduction Checklist
The following sequence reproduces the pipeline against a new source LAS file. Placeholders in angle brackets should be replaced with values specific to the local deployment.
# 1. Create and activate the isolated Entwine environment conda create -n atlas-pointcloud -c conda-forge entwine conda activate atlas-pointcloud # 2. Install the Python dependencies used for publication pip install boto3 laspy # 3. Inspect the source file before conversion pdal info --summary source.las # 4. Build the EPT octree with a restricted schema entwine build \ -i source.las \ -o output/ept \ --dataType binary \ --noOriginId \ --config schema-override.json # 5. Publish the built node tree to object storage behind the CDN aws s3 sync output/ept \ s3://<S3_BUCKET_NAME>/projects/<PROJECT_SLUG>/pointcloud/ept/ # 6. Confirm the published manifest is reachable and CORS-enabled curl -I https://<CLOUDFRONT_DOMAIN>/projects/<PROJECT_SLUG>/pointcloud/ept/ept.json # 7. Run the decode and LOD regression tests against the published dataset node verify-decode.js https://<CLOUDFRONT_DOMAIN>/projects/<PROJECT_SLUG>/pointcloud/ept/ept.json node verify-lod.js https://<CLOUDFRONT_DOMAIN>/projects/<PROJECT_SLUG>/pointcloud/ept/ept.json
