Back to Projects
● ConceptGIS & MappingPoint Cloud Delivery

From LAS to Browser

Building a streaming point-cloud delivery pipeline.

A complete pipeline that converts a raw LAS file into an interactive, browser-rendered point cloud. The pipeline builds an EPT octree, restricts the per-point schema to cut transferred bytes by 2.72x, publishes to object storage behind a CDN, and automates ingestion through a file-stability-aware worker.

LiDAREPTEntwineThree.jsAWS S3CloudFrontPoint CloudWebGL
Read Time~15 min
PublishedAug 2026
StackEntwine + EPT + Three.js + AWS
ATLAS point cloud viewer
01

Overview

Airborne LiDAR surveys routinely produce single files of several gigabytes and tens of millions of points, a volume that cannot be transferred to a web browser as a unit and cannot be rendered by conventional graphics pipelines without discarding the vast majority of the data at any given moment.

The system described here accepts a LAS file as input and produces a tiled representation suitable for range-limited HTTP delivery from a CDN, a browser client that loads a coarse overview immediately and refines it without blocking, and an automated ingestion path that converts a dropped file into a published, browsable dataset without further operator action.

"Streaming," in the sense used throughout this report, does not mean a continuous byte transfer of the whole file. It means restructuring the dataset once, ahead of time, into a hierarchy of spatial tiles ordered from coarse to fine, so that a client can request only the tiles relevant to its current view and refine that view as the camera approaches.

Representative Dataset
For the representative build used throughout this report: 91,625,792 points recorded across a ground footprint of 367 m by 271 m, occupying 2.9 GB in its native LAS format.
02

Background

LAS and LAZ

LAS is a binary file format for storing point-cloud data, standardized by ASPRS. Each point occupies a fixed-size record defined by a "point data record format" (PDRF) number. LAZ is a lossless compressed variant. Neither format is tiled or organized by level of detail: reading any region requires scanning some or all of the file.

The Octree and Level-of-Detail

An octree is a tree structure in which each node represents a cubical region of 3D space, and each non-leaf node's cube is subdivided into eight equal child cubes. Applied to a point cloud, an octree node holds a spatially thinned subset of the points that fall within its cube. The root encloses the entire survey; its children hold finer subsets of smaller sub-volumes.

This structure enables a level-of-detail (LOD) scheme: a renderer displays the root node alone as a coarse representation, then descends into whichever child nodes fall within the camera's view and are worth resolving, replacing coarseness with detail only where a viewer would notice it.

The Entwine Point Tile Format

EPT is a format and directory layout for storing a point cloud as an octree of binary tiles, produced by the Entwine tool. An EPT dataset consists of a manifest file (ept.json) describing bounds, schema, and structure, together with one binary data file per node and a set of hierarchy files describing which nodes exist.

Why EPT?
Because each node is a separate, independently addressable file, a client can fetch exactly the nodes it needs over HTTP from plain static hosting, without server-side logic beyond serving files. The tiling work is done once, at build time, and every subsequent read is a plain file fetch.

The Role of a CDN

A content delivery network caches static files at edge locations distributed relative to end users, so repeated requests for the same file are served from a nearby cache rather than from the origin on every request. For a tiled point cloud, where coarse tiles are requested by every viewer and fine tiles are requested repeatedly by users zooming into the same regions, this cache behaviour keeps load times and origin bandwidth costs bounded as viewership grows.

03

Source Data Characteristics

Before any conversion step, the source LAS file should be inspected to confirm the properties that later pipeline stages assume. Inspection was performed with PDAL, run directly against the source file:

Inspect source file
pdal info --summary source.las

The properties that matter for the stages that follow are the point count and file size, the point data record format, the coordinate reference system, and the spatial extent.

PropertyValue
Point Count91,625,792
File Size2.9 GB
PDRF7 (XYZ, RGB, GPS time)
Ground Footprint367 m × 271 m
Vertical Extent~30 m
04

Environment

Three tools are required before any conversion: Entwine to build the EPT octree, PDAL to inspect and decimate the source file, and a Python 3.13 environment with boto3 and laspy for the publication stage. Entwine was installed in an isolated conda environment, because it links its own build of GDAL and PDAL, and installing it alongside other GIS software risks conflicting shared libraries.

Create isolated environment
conda create -n atlas-pointcloud -c conda-forge entwine
conda activate atlas-pointcloud
pip install boto3 laspy
Windows PATH Pitfall
Entwine, installed inside a conda environment, is not placed on the system PATH. The environment's Library/bin directory must itself be on PATH so that Entwine's own DLLs resolve at runtime. The same class of problem affects gdal2tiles: run from outside its owning environment, it fails with ModuleNotFoundError: No module named '_gdal'. The reported error gives no indication that the actual cause is PATH ordering. Activate the environment and confirm with where entwine (Windows) or which entwine (Unix).
05

Conversion to EPT

With the environment active, the EPT octree is built from the source LAS file:

Build EPT octree
entwine build -i source.las -o output/ept --dataType binary

The --dataType binary flag is not optional for this pipeline. The client renderer decodes raw binary node records directly; it does not implement a LAZ decompressor, and a node written in Entwine's laszip data type is rejected outright. Passing the flag explicitly means a future change to Entwine's default output format cannot silently produce a dataset the existing renderer is unable to read.

For the representative dataset, this build produced 1,914 octree nodes. Each node holds points sampled on a grid of span voxels per edge of its cube, where span is 128.

Output Directory

ept.json

Manifest: bounds, schema, point count

ept-data/

One binary file per octree node

ept-hierarchy/

JSON files describing node existence and tree connections

Hierarchy Continuation Markers
The root hierarchy file lists, for each node, either a positive point count (a leaf) or a negative count (a continuation marker). A negative value means descendants are described in a separate file, ept-hierarchy/<node-id>.json, which the client must fetch in turn. A client that reads only the root file will build a tree that terminates at whatever depth the root describes directly. This produces no error: the symptom is a dataset that renders only its coarsest levels, indistinguishable from one genuinely built to a shallow depth.
06

Payload Optimisation

Entwine's default schema carries 22 dimensions at 49 bytes per point. The renderer consumes only six dimensions (X, Y, Z, Red, Green, Blue), meaning 62 percent of every transferred byte is discarded on arrival. The schema was restricted to these six dimensions:

Build with restricted schema
entwine build \
  -i source.las \
  -o output/ept \
  --dataType binary \
  --noOriginId \
  --config schema-override.json
schema-override.json
{
  "schema": [
    { "name": "X" },
    { "name": "Y" },
    { "name": "Z" },
    { "name": "Red" },
    { "name": "Green" },
    { "name": "Blue" }
  ]
}
Omit the Scale Key Deliberately
The dimension entries specify only a name, with no scale key. Entwine derives an appropriate coordinate scale from the source point cloud's extent and precision when none is given. Supplying a fixed scale value would silently coarsen coordinate precision regardless of whether that value suited the source data, producing geometrically incorrect output with no error raised at build time. The scale key should be absent, not merely left at a default value.

Restricting the schema reduced the per-point record from 49 to 18 bytes, a factor of 2.72. Decoded geometry and colour from the restricted build were verified byte-identical to the corresponding values from the unrestricted build.

MetricDefault SchemaRestricted SchemaReduction
Bytes per point49 B18 B2.72x
Total node data4,194 MB~1,541 MB2.72x
Median node file~2 MB~760 kB2.63x
Why Not Just Compress?
Compression of already-binary, high-entropy point data yields a materially smaller reduction than simply not transferring unused fields. It also does nothing to reduce the on-disk storage footprint or the number of bytes the client must still parse after decompression.
07

Publication

The built node tree is synchronized to object storage under a per-project prefix and served through a CDN:

Sync to S3
aws s3 sync output/ept \
  s3://<S3_BUCKET_NAME>/projects/<PROJECT_SLUG>/pointcloud/ept/
CORS Is a Functional Requirement
The Access-Control-Allow-Origin response header must be present on responses for ept.json, the hierarchy files, and the node data files, or the browser's fetch requests are blocked by the browser itself regardless of whether the underlying storage request would have succeeded. A dataset published without this header appears present and correctly formed from the operator's perspective, while remaining entirely unreachable from the client renderer.

Publication emits a machine-readable summary containing the public URL of the published ept.json, which is written back to the corresponding project's database record so the client application knows where to fetch the dataset without further manual configuration.

08

Automation

A manual build-and-publish procedure does not scale past the first survey. The pipeline is triggered by a local worker process that watches an input directory; dropping a file named pointcloud.las into a project's folder enqueues a build job.

File-Stability Aware Ingestion
A multi-gigabyte copy operation creates the destination filename at the instant the copy begins, not when it finishes. A naive watcher that acts on first appearance hands the builder a truncated, partially written file. The worker requires three conditions before enqueuing: the file's size and modification time must be stable across several consecutive polls, and the file must be openable for writing by the worker (a check that fails while the copying process still holds the write handle).

Completed work is recorded as a fingerprint of the file's size and modification time at the moment it was processed, so restarting the worker does not reprocess every file already present. Fingerprints are recorded for failed builds as well as successful ones; without this, a file that reliably crashes the builder would be retried indefinitely on every restart rather than surfaced once as a failure requiring operator attention.

09

Client Architecture

The browser-side renderer is built on three.js, using a single THREE.Points object per loaded octree node and one shared PointsMaterial across all of them. A custom scheduler fetches the manifest, walks the full hierarchy including continuation files, renders the root node immediately so a coarse view appears without waiting for deeper fetches, then progressively requests finer nodes as the camera and the loading gate permit.

PositionFloat32Array

World-space coordinates require float32 precision

ColourUint8Array (normalized)

256 levels per channel is sufficient; GPU normalizes to 0-1

This reduces the colour attribute from 12 bytes per point (three 32-bit floats) to 3 bytes, and the combined per-point GPU footprint from 24 to 15 bytes, a reduction of 37.5 percent, while the vertex shader still receives colour values in the 0-to-1 range it expects.

10

Decoding EPT Coordinates

EPT stores each point's coordinates not as raw floating-point positions but as scaled integers: the true coordinate is raw × scale + offset, where scale and offset are given per dimension in the schema.

The Defect That Threw No Error
An initial implementation omitted the transform entirely and mapped the raw int32 range directly across each node's bounding cube. Because real coordinate values are tiny relative to the full range of a 32-bit signed integer, every point within a node collapsed onto a single position. A true spread of 285.77 m rendered as a spread of 2.4 mm. The symptom was a sparse scatter of roughly one visible dot per loaded node, not an obviously broken scene. No error was thrown at any stage, because a coordinate transform applied incorrectly still produces a buffer of finite, well-formed floating-point numbers.
Bounds vs BoundsConforming
The bounds field in ept.json describes the octree's enclosing cube, which is necessarily larger than the surveyed data. A separate field, boundsConforming, gives the true, non-cubical extent of the data. Centring the scene on the enclosing cube rather than on boundsConforming placed roughly half of the point cloud below the intended reference grid.
11

Performance Engineering

The governing principle: a point the screen cannot resolve should never be retained in memory or on the GPU. Each node samples its cube on a grid with span divisions per edge, so adjacent points sit edge / span apart in world units. Projecting this spacing to screen pixels, given the node's distance from the camera, tests directly whether loading that node adds visible detail.

Division of Labour
A resolution-based gate governs what is worth loading. A budget-based evictor governs how much of what qualifies is retained. At a wide viewing distance, the pixel-spacing gate binds: the resident set measured 3.07 M points identically at both a 6 M and 12 M budget tier. At close range, the budget itself binds. This is the intended behaviour, not an inconsistency.

Three Corrected Defects

  • 01

    No threshold on loading

    The scheduler queued every node in the frustum with no threshold on whether loading it would produce a visible change. Each camera movement pulled in strictly more data than the last, until the entire dataset was resident. The symptom was deceptive: performance was acceptable for the first few seconds and degraded progressively thereafter.

  • 02

    Eviction considered only out-of-frustum nodes

    Because nearly every loaded node remains inside the frustum at any given moment, this left no eviction candidates, meaning the point budget could not be enforced at all. It also rendered the adaptive frame-rate tuner inert.

  • 03

    LRU timestamp never updated

    Each node's lastAccess timestamp was written once, at insertion, and never updated on subsequent frames. LRU ordering was equivalent to plain insertion order, so eviction removed nodes unrelated to how recently they had been rendered.

12

Verification Methodology

Two standalone Node.js scripts assert directly against the live, published dataset rather than against local fixtures, since the defects described above are properties of the full-scale, deployed data and its network delivery.

verify-decode.jsDecoding defect

Asserts on the spread and bounds of decoded point coordinates, not merely the type or shape of the buffer

verify-lod.jsLoading & eviction

Drives 60 simulated camera movements and asserts the resident count converges rather than climbing without bound

Static Checks Are Not Enough
During refactoring, two internal methods were accidentally deleted in a way that left their call sites intact. Linting and the production build both completed without error. The decode regression test, which asserts on numeric properties of decoded output, caught the regression. Static checks verify only that code is well-formed; they cannot verify that a computation still produces the numeric result it is supposed to produce.
13

Results

The pipeline transforms a 2.9 GB LAS file into a CDN-served EPT dataset that a browser can render interactively. The schema restriction cuts transferred bytes by 2.72x, the client keeps the resident point count between one and two orders of magnitude below the source dataset, and the pixel-resolution loading gate ensures a point the screen cannot resolve is never retained.

MetricBeforeAfter
Per-point payload49 B18 B
Total node data4,194 MB~1,541 MB
Resident points (wide view)Unbounded3.07 M pts
GPU per-point footprint24 B15 B
14

Limitations and Future Work

  • No compression

    The --dataType binary requirement forgoes Entwine's laszip output. A compressed format compatible with client-side decoding is a possible further avenue.

  • Single-worker automation

    The automation is a single local worker watching one input path per project. Its behaviour under simultaneous uploads across projects, or recovery after a crash mid-build, is not yet covered.

  • Platform-specific PATH issues

    The Windows-specific PATH pitfalls were observed on that platform; equivalent pitfalls on macOS or Linux are not addressed.

  • Fixed device-pixel-ratio cap

    The device-pixel-ratio cap is a fixed 1.5 chosen for the hardware on which it was measured. An adaptive cap responding to a monitored frame budget is a plausible extension.

15

Reproduction Checklist

The following sequence reproduces the pipeline against a new source LAS file. Placeholders in angle brackets should be replaced with values specific to the local deployment.

Reproduction
# 1. Create and activate the isolated Entwine environment
conda create -n atlas-pointcloud -c conda-forge entwine
conda activate atlas-pointcloud

# 2. Install the Python dependencies used for publication
pip install boto3 laspy

# 3. Inspect the source file before conversion
pdal info --summary source.las

# 4. Build the EPT octree with a restricted schema
entwine build \
  -i source.las \
  -o output/ept \
  --dataType binary \
  --noOriginId \
  --config schema-override.json

# 5. Publish the built node tree to object storage behind the CDN
aws s3 sync output/ept \
  s3://<S3_BUCKET_NAME>/projects/<PROJECT_SLUG>/pointcloud/ept/

# 6. Confirm the published manifest is reachable and CORS-enabled
curl -I https://<CLOUDFRONT_DOMAIN>/projects/<PROJECT_SLUG>/pointcloud/ept/ept.json

# 7. Run the decode and LOD regression tests against the published dataset
node verify-decode.js https://<CLOUDFRONT_DOMAIN>/projects/<PROJECT_SLUG>/pointcloud/ept/ept.json
node verify-lod.js    https://<CLOUDFRONT_DOMAIN>/projects/<PROJECT_SLUG>/pointcloud/ept/ept.json