Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,31 @@ On read: a PLY header comment — Brush's `comment SplatRenderMode: default | mi

On write: `.ply` and `.compressed.ply` carry `comment SplatRenderMode: mip | 2dgs` (Brush's spelling, whichever form was read); `.sog` and `meta.json` carry `"model": "antialiased" | "2dgs"`; `.spz` sets its antialiased bit, and warns that it cannot represent 2DGS. A 2DGS PLY output drops the `scale_2` column again. Other output formats have nowhere to record it and drop the tag silently. Combining inputs whose models disagree warns and writes the result untagged.

### Scene camera

A SOG can carry one optional camera in `meta.json` — a rest pose, the pinhole intrinsics and a few viewing hints — so a viewer can open the scene where it was meant to be seen without a sidecar file. It matters most for splats lifted from a single photo or stereo pair, which only look right from near the camera that took it. The block is advisory and readers that don't use it ignore it:

```json
"camera": {
"convention": "opencv",
"rig": "camera",
"rest": { "position": [0, 0, 0], "rotation": [0, 0, 0, 1] },
"intrinsics": { "fx": 1194.67, "fy": 1194.67, "cx": 1024, "cy": 576, "width": 2048, "height": 1152 },
"stereo": { "baseline_m": 0.063 },
"focus": { "point": [0, 0, 1.68], "subject_m": 2.14, "near_m": 0.73, "far_m": 66.2 }
}
```

`rest.position` and `rest.rotation` (camera-to-world, `[x, y, z, w]`) are in the same coordinates as the gaussians; `convention` names the camera axes (`opencv`: +x right, +y down, +z forward); `intrinsics` are in pixels of a `width` × `height` image. Everything is optional.

On read, a `.sog`, `meta.json` or `lod-meta.json` camera is kept; on write, `.sog` and `meta.json` carry it in `meta.json`, and `lod-meta.json` carries it once at its top level. Translate, rotate and scale actions move the rest pose with the scene (and scale the `focus` and `stereo` distances); other keys in the block are passed through unchanged. Other output formats have nowhere to record it and drop it. When several inputs carry a camera, the first is kept.

`--camera-from cameras.json[:n]` sets it from a training camera in the `cameras.json` that 3DGS trainers write next to the PLY (`position`, `rotation`, `fx`, `fy`, `width`, `height`, OpenCV axes; the principal point is taken as the image centre):

```bash
splat-transform scene.ply --camera-from cameras.json:12 scene.sog
```

## Actions

Actions execute in the order specified and can be repeated. Any action may appear after any input or output file:
Expand Down Expand Up @@ -142,6 +167,9 @@ Actions execute in the order specified and can be repeated. Any action may appea
--stats [text|json] Print file info, per-column statistics and the fill/overdraw ratio to stdout. Default: text
--info [text|json] Print structural metadata (format, per-LOD counts, extra columns) to stdout. Default: text
-m, --morton-order Reorder Gaussians by Morton code (Z-order curve)
--camera-from <file[:n]> Record training camera n (default 0) of a 3DGS cameras.json
as the input's camera (see Scene camera). Place it after
the input the poses belong to.
```

## CLI Options
Expand Down
40 changes: 35 additions & 5 deletions src/cli/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ import {
resolveSplatModel,
revision,
selectLod,
sogCameraFromCamerasJson,
stackLods,
TextRenderer,
Transform,
Expand All @@ -37,6 +38,7 @@ import {
WorkerQueue,
writeLodSource,
writeSource,
withCamera,
type ChunkSource,
type ChunkSourceMetadata,
type ProcessAction,
Expand All @@ -45,6 +47,7 @@ import {
type Options as LibOptions,
type CollisionMeshShape,
type ReadFileSystem,
type SogCamera,
logger
} from '../lib';
// CLI-only internals (deliberately off the public lib surface): the LOD-path
Expand Down Expand Up @@ -144,6 +147,8 @@ const stripLodTags = (actions: CliAction[]): ProcessAction[] => {
type File = {
filename: string;
processActions: CliAction[];
/** Camera block from `--camera-from`, in this input's own coordinates. */
camera?: SogCamera;
};

const cliOptionsConfig = {
Expand Down Expand Up @@ -214,7 +219,8 @@ const cliOptionsConfig = {
'tag-lod': { type: 'string', short: 'l', multiple: true },
stats: { type: 'string', multiple: true },
info: { type: 'string', multiple: true },
'morton-order': { type: 'boolean', short: 'm', multiple: true }
'morton-order': { type: 'boolean', short: 'm', multiple: true },
'camera-from': { type: 'string', multiple: true }
} as const;

const stringOptionNames = new Set(Object.entries(cliOptionsConfig)
Expand Down Expand Up @@ -822,6 +828,19 @@ const parseArguments = async () => {
current.processActions.push(ffAction);
break;
}
case 'camera-from': {
// <cameras.json>[:index], index defaulting to the first camera
const match = /^(.*):(\d+)$/.exec(t.value);
const path = match ? match[1] : t.value;
let cameras;
try {
cameras = JSON.parse(await pathReadFile(path, 'utf-8'));
} catch (e) {
throw new Error(`Failed to read cameras JSON file: ${path}`);
}
current.camera = sogCameraFromCamerasJson(cameras, match ? parseInteger(match[2]) : 0);
break;
}
}
}
}
Expand Down Expand Up @@ -874,6 +893,8 @@ ACTIONS (executed in order; can be repeated)
--stats [text|json] Print file info, per-column statistics and the fill/overdraw ratio to stdout. Default: text
--info [text|json] Print structural metadata (format, per-LOD counts, extra columns) to stdout. Default: text
-m, --morton-order Reorder Gaussians by Morton code (Z-order curve)
--camera-from <file[:n]> Record training camera n (default 0) of a 3DGS cameras.json as the
input's camera; written to .sog / meta.json / lod-meta.json

GENERAL
-h, --help Show this help and exit
Expand Down Expand Up @@ -1137,6 +1158,10 @@ const main = async () => {

const outputFilename = resolve(outputArg.filename);

if (outputArg.camera) {
failExit('--camera-from applies to an input file: place it after the input whose poses it holds.');
}

// Check for null output (discard file writing)
const isNullOutput = outputArg.filename.toLowerCase() === 'null';

Expand Down Expand Up @@ -1289,7 +1314,8 @@ const main = async () => {
});
const readFilename = fmt === 'mjs' ? `file://${inFile}` : inFile;
const srcs = await readFile({ filename: readFilename, inputFormat: fmt, options: { ...options, lodSelect: [] }, params, fileSystem });
return srcs.length === 1 ? srcs[0] : concatSource(srcs, pool);
const src = srcs.length === 1 ? srcs[0] : concatSource(srcs, pool);
return inputArg.camera ? withCamera(src, inputArg.camera) : src;
};

// Stitch inputs: uniform layout -> concatSource (transforms unified as
Expand All @@ -1312,12 +1338,14 @@ const main = async () => {
const seen = [...new Set(sources.map(s => s.meta.model))].join(', ');
logger.warn(`mixed splat models (${seen}); writing the result as '${model}'`);
}
const withCam = sources.find(s => s.meta.camera);
const dts: DataTable[] = [];
for (const s of sources) {
dts.push(await materializeToDataTable(s, pool));
await s.close();
}
return dataTableToChunkSource(combine(dts), pool.chunkSize, undefined, model);
const combinedSource = dataTableToChunkSource(combine(dts), pool.chunkSize, undefined, model);
return withCam ? withCamera(combinedSource, withCam.meta.camera, withCam.meta.transform) : combinedSource;
};

const phase = logger.group(`Output ${outputArg.filename}`, { index: phaseTotal, total: phaseTotal });
Expand Down Expand Up @@ -1416,11 +1444,12 @@ const main = async () => {
// Intrinsic multi-LOD: view each level with selectLod (shared parent);
// env fetched separately. The input's own actions apply per level.
const { filename: inFile, fileSystem } = resolveInput(inputArgs[0].filename);
const multi = single === 'lcc2' ?
const lodSource = single === 'lcc2' ?
await readLcc2Source(fileSystem, inFile, { ...options, lodSelect: [] }, pool) :
single === 'lod' ?
await readLodSource(fileSystem, inFile, { ...options, lodSelect: [] }, pool) :
await readLccSource(fileSystem, inFile, { ...options, lodSelect: [] }, pool);
const multi = inputArgs[0].camera ? withCamera(lodSource, inputArgs[0].camera) : lodSource;
container = multi;
envSource = single === 'lcc2' ?
await readLcc2EnvironmentSource(fileSystem, inFile, pool) :
Expand All @@ -1445,8 +1474,9 @@ const main = async () => {
}
const opened = await Promise.all(tagged.map(async (t) => {
const { filename: inFile, fileSystem } = resolveInput(t.arg.filename);
const ply = await readPly(await fileSystem.createSource(inFile), pool);
const src = await processSourceBridged(
await readPly(await fileSystem.createSource(inFile), pool),
t.arg.camera ? withCamera(ply, t.arg.camera) : ply,
t.rest, pool, processOptions
);
return { src, tag: t.tag };
Expand Down
6 changes: 6 additions & 0 deletions src/lib/chunk/source.ts
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
import { type SogCamera } from '../sog-camera';
import { type SplatModel } from '../splat-model';
import { type Transform } from '../utils';
import { type ChunkData } from './data';
Expand Down Expand Up @@ -38,6 +39,11 @@ type ChunkSourceMetadata = {
readonly availableLayers: ReadonlySet<ChunkLayer>;
/** Per-layer stride + field map. Keyed by layer; only present for available layers. */
readonly layouts: Readonly<Partial<Record<ChunkLayer, LayerLayout>>>;
/**
* Optional capture / rest camera (SOG `meta.json` `camera`). Like the data,
* it is stored raw: `transform` gives its meaning, and baking moves it along.
*/
readonly camera?: SogCamera;
};

/**
Expand Down
1 change: 1 addition & 0 deletions src/lib/decimate-uniform/decimate-source.ts
Original file line number Diff line number Diff line change
Expand Up @@ -250,6 +250,7 @@ const decimateSource = async (
numChunks: [Math.ceil(outCount / src.meta.chunkSize)],
shBands: src.meta.shBands,
model: src.meta.model,
camera: src.meta.camera,
extraColumns: src.meta.extraColumns,
transform: src.meta.transform,
availableLayers: src.meta.availableLayers,
Expand Down
1 change: 1 addition & 0 deletions src/lib/decimate/decimate-source.ts
Original file line number Diff line number Diff line change
Expand Up @@ -433,6 +433,7 @@ const decimateSource = async (
numChunks: [Math.ceil(outCount / src.meta.chunkSize)],
shBands: src.meta.shBands,
model: src.meta.model,
camera: src.meta.camera,
extraColumns: src.meta.extraColumns,
transform: src.meta.transform,
availableLayers: src.meta.availableLayers,
Expand Down
7 changes: 6 additions & 1 deletion src/lib/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -12,13 +12,18 @@ export type {
} from './chunk';

// Structural combinators (lazy views over sources)
export { bakeTransform, concatSource, selectLod, stackLods, sortMortonColumns, sortMortonInterleaved } from './ops';
export { bakeTransform, concatSource, selectLod, stackLods, sortMortonColumns, sortMortonInterleaved, withCamera } from './ops';

// How a scene was trained (carried on `ChunkSourceMetadata.model`). The per-format
// spellings of the tag live with their reader/writer.
export { isSplatModel, resolveSplatModel } from './splat-model';
export type { SplatModel } from './splat-model';

// The optional camera block of SOG meta.json / lod-meta.json (carried on
// `ChunkSourceMetadata.camera`).
export { sogCameraFromCamerasJson, transformSogCamera } from './sog-camera';
export type { SogCamera } from './sog-camera';

// Action processing over a source: `processSource` streams and throws on
// actions that need the DataTable bridge; `processSourceBridged` handles every
// action, materializing only the DataTable-only runs as islands.
Expand Down
9 changes: 8 additions & 1 deletion src/lib/ops/bake-transform.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ import {
type ChunkSourceMetadata,
SH_REST_COUNTS
} from '../chunk';
import { transformSogCamera } from '../sog-camera';
import { RotateSH, Transform } from '../utils';

const SH_PER_CHANNEL = [0, 3, 8, 15];
Expand All @@ -33,13 +34,19 @@ const SH_PER_CHANNEL = [0, 3, 8, 15];
* @returns A derived source whose reads yield data in `targetSpace`.
*/
const bakeTransform = (src: ChunkSource, targetSpace: Transform): ChunkSource => {
const meta: ChunkSourceMetadata = { ...src.meta, transform: targetSpace.clone() };
const delta = targetSpace.clone().invert().mul(src.meta.transform);

if (delta.isIdentity()) {
const meta: ChunkSourceMetadata = { ...src.meta, transform: targetSpace.clone() };
return { meta, read: req => src.read(req), close: () => src.close() };
}

const meta: ChunkSourceMetadata = {
...src.meta,
transform: targetSpace.clone(),
...(src.meta.camera ? { camera: transformSogCamera(src.meta.camera, delta) } : {})
};

const r = delta.rotation;
const s = delta.scale;
const rx = r.x, ry = r.y, rz = r.z, rw = r.w;
Expand Down
9 changes: 9 additions & 0 deletions src/lib/ops/concat-source.ts
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,14 @@ const concatSource = (allSources: ChunkSource[], pool: ChunkDataPool): ChunkSour
logger.warn(`mixed splat models (${seen}); writing the result as '${model}'`);
}

// One output holds one camera: keep the first input's (the inputs share a
// transform, so it needs no re-expressing), and say so if another disagreed.
const cameras = allSources.map(s => s.meta.camera).filter(c => c !== undefined);
const camera = cameras[0];
if (cameras.some(c => JSON.stringify(c) !== JSON.stringify(camera))) {
logger.warn('inputs carry different cameras; keeping the first');
}

const S = ref.chunkSize;
// Per-source gaussian counts and the output-row offset each source begins at.
const counts = sources.map(s => s.meta.numGaussians);
Expand All @@ -100,6 +108,7 @@ const concatSource = (allSources: ChunkSource[], pool: ChunkDataPool): ChunkSour
const meta: ChunkSourceMetadata = {
...ref,
model,
camera,
numGaussians: total,
numLods: 1,
lodCounts: [total],
Expand Down
1 change: 1 addition & 0 deletions src/lib/ops/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ export { selectLod, resolveLodLevels } from './select-lod';
export { filterSource } from './filter-source';
export { reduceBandsSource } from './reduce-bands-source';
export { concatSource } from './concat-source';
export { withCamera } from './with-camera';
export { filterNaNRows, filterByValueRows, filterBoxRows, filterSphereRows } from './filter-mask';
export { computeSourceStats } from './stats';
export type { LodStats, LodStatsData, SourceStats } from './stats';
32 changes: 32 additions & 0 deletions src/lib/ops/with-camera.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
import { type ChunkSource, type ChunkSourceMetadata } from '../chunk';
import { type SogCamera, transformSogCamera } from '../sog-camera';
import { type Transform } from '../utils';

/**
* Set (or, with `undefined`, clear) a source's camera block, lazily. Reads pass
* through unchanged.
*
* `space` is the pending transform the camera's values were stored under — by
* default the source's own, i.e. the camera is in the same raw coordinates as
* the source's gaussians. When it differs (the source was re-bridged or baked
* since), the camera is re-expressed so that it still means the same pose.
*
* @param src - The parent source.
* @param camera - The camera block, or `undefined` to drop it.
* @param space - The pending transform `camera` is stored under. Default: `src.meta.transform`.
* @returns A derived source carrying the camera.
*/
const withCamera = (src: ChunkSource, camera: SogCamera | undefined, space?: Transform): ChunkSource => {
const { transform } = src.meta;
const cam = camera && space && !space.equals(transform) ?
transformSogCamera(camera, transform.clone().invert().mul(space)) :
camera;
const meta: ChunkSourceMetadata = { ...src.meta, camera: cam };
return {
meta,
read: req => src.read(req),
close: () => src.close()
};
};

export { withCamera };
5 changes: 4 additions & 1 deletion src/lib/process-source.ts
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,8 @@ import {
mapSource,
mortonOrder,
permuteSource,
reduceBandsSource
reduceBandsSource,
withCamera
} from './ops';
import { processDataTable, type ProcessAction, type ProcessOptions } from './process';
import { formatSourceInfo, formatSourceStats } from './source-info';
Expand Down Expand Up @@ -151,9 +152,11 @@ const processSourceBridged = async (
} else {
// DataTable island: materialize the current (streaming) source, apply
// the run on the table, and re-bridge back to a source to keep going.
const { camera, transform } = src.meta;
const dt = await materializeToDataTable(src, pool);
await src.close();
src = dataTableToChunkSource(await processDataTable(dt, run, options), pool.chunkSize);
if (camera) src = withCamera(src, camera, transform);
}
i = j;
}
Expand Down
7 changes: 6 additions & 1 deletion src/lib/readers/read-lod.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@ import { containerSource, type ContainerSegment } from './container-source';
import { readSogSource } from './read-sog';
import { type ChunkDataPool, type ChunkSource } from '../chunk';
import { dirname, join, readFile, type ReadFileSystem } from '../io/read';
import { withCamera } from '../ops';
import { readSogCamera, type SogCamera } from '../sog-camera';
import { type Options } from '../types';

type LodReference = {
Expand All @@ -22,6 +24,7 @@ type LodMeta = {
counts?: number[];
lodLevels: number;
environment?: string;
camera?: SogCamera;
filenames: string[];
tree: LodNode;
};
Expand Down Expand Up @@ -155,7 +158,9 @@ const readLodSource = async (
});
});

return containerSource(segmentsByLod, pool);
const source = await containerSource(segmentsByLod, pool);
const camera = readSogCamera(meta.camera);
return camera ? withCamera(source, camera) : source;
};

/**
Expand Down
Loading