From 60ff7004cadfd6ffb138fddf2dcbc996ba641ced Mon Sep 17 00:00:00 2001 From: thanos Date: Fri, 17 Jul 2026 17:26:14 -0400 Subject: [PATCH 1/6] closed #37 - `EXCP` and `GSPL` file stream_decode now do real incremental I/O. MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit **Behavior** `source: :file` (or `:auto` path detection): read 16-byte EXCP / 18-byte GSPL header, then one fixed-size record via `IO.binread/2` Peak memory ≈ header + one record (file handle stays open for the stream lifetime) In-memory binaries still materialize through `decode/2` (mmap left for a future Rust backend) **Wiring** `ExCodecs.Spatial.Stream` forwards `EXCP/GSPL` sources to the codecs — no more full File.read for those formats Docs updated in `spatial.ex`, `docs/spatial_formats.md`, the spatial guide, and `CHANGELOG` Unreleased **Usage** ```elixir ExCodecs.Spatial.stream_decode(path, format: :spatial_binary, source: :file) |> Stream.take(100) |> Enum.to_list() ExCodecs.Spatial.stream_decode(path, format: :gsplat, source: :file) |> Enum.each(&process/1) ``` Truncation / missing files yield a single `{:error, %ExCodecs.Error{}}` element. closed #38 - PLY stream_decode (source: :file) — scans until end_header, then yields one vertex at a time (binary: fixed stride via IO.binread; ASCII: line reads). Supports :as / Gaussian auto-detect. closed #39 - stream_encode_to_file/3 — EXCP and GSPL write a placeholder header, stream records, seek back to patch count. Requires explicit :schema. Spatial.Stream.encode_to_file/3 uses this when schema: + format: :spatial_binary / :gsplat. closed #40 --- CHANGELOG.md | 17 + docs/spatial_formats.md | 24 +- guides/understanding_spatial_codecs.md | 5 +- lib/ex_codecs/spatial.ex | 5 +- lib/ex_codecs/spatial/codec/binary.ex | 362 +++++++++++++++++-- lib/ex_codecs/spatial/codec/gsplat.ex | 325 +++++++++++++++-- lib/ex_codecs/spatial/codec/ply.ex | 297 ++++++++++++--- lib/ex_codecs/spatial/stream.ex | 141 ++++---- test/ex_codecs/spatial/codec/binary_test.exs | 77 ++++ test/ex_codecs/spatial/codec/gsplat_test.exs | 71 ++++ test/ex_codecs/spatial/codec/ply_test.exs | 85 +++++ test/ex_codecs/spatial/stream_test.exs | 33 ++ 12 files changed, 1280 insertions(+), 162 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6727bbb..a76a3a2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,23 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Changed + +- EXCP (`:spatial_binary`), GSPL (`:gsplat`), and PLY `stream_decode` with + `source: :file` (or `:auto` path detection) now read the header and then + **one record/vertex at a time** from disk (bounded memory). Binary PLY uses + a fixed stride; ASCII PLY reads lines. In-memory binaries still materialize; + a future Rust backend may memory-map those. + +### Added + +- `Binary.stream_encode_to_file/3` and `Gsplat.stream_encode_to_file/3` — + incremental file writes with an explicit `:schema` (placeholder header, + seek-back count). `Spatial.Stream.encode_to_file/3` uses these when + `:schema` is present with `format: :spatial_binary` or `:gsplat`. + ## [0.2.0] - 2026-07-16 ### Added diff --git a/docs/spatial_formats.md b/docs/spatial_formats.md index 4fc83f5..d1beb1f 100644 --- a/docs/spatial_formats.md +++ b/docs/spatial_formats.md @@ -42,9 +42,16 @@ normal `{0,0,0}`). Mixed optional fields therefore round-trip with defaults fill ### Streaming -`stream_decode` / `stream_encode` currently **materialize** the full payload (or -enumerable) then enumerate. Prefer explicit `source: :file` or `source: :binary` -when the argument is ambiguous. +- **EXCP / GSPL / PLY + `source: :file`** (or `:auto` path detection): header + then one record/vertex at a time from disk (bounded memory). Binary PLY uses + a fixed property stride; ASCII PLY reads lines. +- In-memory binaries still **materialize** through `decode/2`, then enumerate. +- `stream_encode` still collects the enumerable, then encodes once. +- `encode_to_file/3` with an explicit `:schema` streams EXCP/GSPL to disk + (placeholder header + seek-back count). Without `:schema`, encode then write. + +Prefer explicit `source: :file` or `source: :binary` when the argument is +ambiguous. A future Rust backend may memory-map large in-memory binaries. `:auto` treats a binary as a path only when it looks path-like (under 4 KiB, no `ply`/`EXCP`/`GSPL` magic, and contains `/` or `\` or ends with `.ply`/`.excp`/ @@ -82,6 +89,11 @@ zero-filled (alpha default 255). after `count` records are ignored. Truncation yields `:invalid_data` / `:truncated_input` as implemented. +**Streaming:** `stream_decode` with `source: :file` reads the 16-byte header, +then each record via `IO.binread/2` (bounded memory). In-memory binaries still +materialize through `decode/2`. `stream_encode_to_file/3` requires `:schema` +(e.g. `schema: [:color]`) and patches the count after writing records. + ## GSPL — `:gsplat` (version 1) Little-endian compact Gaussian clouds. @@ -110,6 +122,12 @@ with zeros on encode). **Not stored:** per-Gaussian metadata maps, cloud metadata. Flags on decode are currently informational; `sh_rest` count in the header is authoritative. +**Streaming:** `stream_decode` with `source: :file` reads the 18-byte header, +then each record via `IO.binread/2` (bounded memory). In-memory binaries still +materialize through `decode/2`. `stream_encode_to_file/3` requires `:schema` +(e.g. `schema: []` or `schema: [sh_rest: 6]`) and patches the count after +writing records. + ## Integrity v1 formats have **no checksum**. For integrity, wrap payloads with a registry diff --git a/guides/understanding_spatial_codecs.md b/guides/understanding_spatial_codecs.md index 56843b2..ea09324 100644 --- a/guides/understanding_spatial_codecs.md +++ b/guides/understanding_spatial_codecs.md @@ -368,8 +368,9 @@ clouds. Magic bytes: `"GSPL"`. Wire layouts and schema rules are frozen in [Spatial wire formats](../docs/spatial_formats.md). -Stream helpers today **materialize** the full payload, then enumerate. Prefer -explicit `source: :file` or `source: :binary`. +Stream helpers: EXCP/GSPL **files** decode record-by-record from disk; PLY and +in-memory binaries still materialize. Prefer explicit `source: :file` or +`source: :binary`. ## Why a separate `ExCodecs.Spatial` API? diff --git a/lib/ex_codecs/spatial.ex b/lib/ex_codecs/spatial.ex index b00a12f..640ccce 100644 --- a/lib/ex_codecs/spatial.ex +++ b/lib/ex_codecs/spatial.ex @@ -46,8 +46,9 @@ defmodule ExCodecs.Spatial do ## Streaming note - `stream_decode` / `stream_encode` currently materialize full payloads, then - enumerate. Prefer `source: :file` when the argument is a path, or + EXCP (`:spatial_binary`) and GSPL (`:gsplat`) **file** sources stream + record-by-record from disk. PLY and in-memory binaries still materialize, + then enumerate. Prefer `source: :file` for large EXCP/GSPL paths, or `source: :binary` for payloads. See `docs/spatial_formats.md` for `:auto` path heuristics and wire-format layouts. """ diff --git a/lib/ex_codecs/spatial/codec/binary.ex b/lib/ex_codecs/spatial/codec/binary.ex index 5fabef6..b7c3457 100644 --- a/lib/ex_codecs/spatial/codec/binary.ex +++ b/lib/ex_codecs/spatial/codec/binary.ex @@ -105,6 +105,148 @@ defmodule ExCodecs.Spatial.Codec.Binary do )} end + @doc """ + Streams `%Point{}` values to an EXCP file using an explicit schema. + + Writes a placeholder header, encodes each point as it arrives, then seeks + back to patch the final count. Peak memory is O(one point), not the cloud. + + ## Arguments + + * `enumerable` (`Enumerable.t()`) — `%Point{}` elements. + * `path` (`Path.t()`) — destination file path. + * `opts` (`keyword()`) — requires `:schema`, a list such as `[]`, + `[:color]`, `[:color, :alpha]`, `[:normal]`, or `[:color, :normal]`. + Map form `%{color: true, alpha: false, normal: true}` is also accepted. + `:alpha` implies color bytes (RGBA). + + ## Returns + + * `:ok` + * `{:error, %ExCodecs.Error{reason: :invalid_options}}` when `:schema` is + missing or invalid + * `{:error, %ExCodecs.Error{reason: :invalid_data}}` when an element is not + a `%Point{}` + * `{:error, %ExCodecs.Error{reason: :io_error}}` on file failures + + ## Examples + + iex> alias ExCodecs.Spatial.{Point, Codec.Binary} + iex> path = Path.join(System.tmp_dir!(), "excp_enc_#{System.unique_integer([:positive])}.excp") + iex> :ok = Binary.stream_encode_to_file([Point.new(1, 2, 3, color: {1, 2, 3})], path, schema: [:color]) + iex> {:ok, <<"EXCP", _::binary>>} = File.read(path) + iex> File.rm!(path) + :ok + """ + @spec stream_encode_to_file(Enumerable.t(), Path.t(), keyword()) :: :ok | {:error, Error.t()} + def stream_encode_to_file(enumerable, path, opts \\ []) do + with {:ok, flags} <- fetch_schema_flags(opts), + {:ok, io} <- open_write(path) do + try do + :ok = IO.binwrite(io, excp_header(flags, 0)) + + count = + Enum.reduce(enumerable, 0, fn + %Point{} = point, n -> + :ok = IO.binwrite(io, encode_point(point, flags)) + n + 1 + + other, _n -> + throw({:bad_point, other}) + end) + + {:ok, 0} = :file.position(io, 0) + :ok = IO.binwrite(io, excp_header(flags, count)) + :ok + catch + {:bad_point, other} -> + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "EXCP stream encode expects Point structs, got: #{inspect(other)}" + )} + after + File.close(io) + end + end + end + + defp fetch_schema_flags(opts) do + case Keyword.fetch(opts, :schema) do + :error -> + {:error, + Error.new(:invalid_options, + codec: :spatial_binary, + message: "EXCP stream_encode_to_file requires schema: (e.g. schema: [:color])" + )} + + {:ok, schema} -> + schema_to_flags(schema) + end + end + + defp schema_to_flags(schema) when is_list(schema) do + alpha? = schema_flag?(schema, :alpha) + color? = alpha? or schema_flag?(schema, :color) + normal? = schema_flag?(schema, :normal) + + unknown = + schema + |> Enum.reject(fn + a when is_atom(a) -> a in [:color, :alpha, :normal] + {k, _} when is_atom(k) -> k in [:color, :alpha, :normal] + _ -> false + end) + + if unknown != [] do + {:error, + Error.new(:invalid_options, + codec: :spatial_binary, + message: "Unknown EXCP schema entries: #{inspect(unknown)}" + )} + else + {:ok, flags_from_schema(color?, alpha?, normal?)} + end + end + + defp schema_to_flags(schema) when is_map(schema) do + schema_to_flags(Map.to_list(schema)) + end + + defp schema_to_flags(other) do + {:error, + Error.new(:invalid_options, + codec: :spatial_binary, + message: "EXCP schema must be a list or map, got: #{inspect(other)}" + )} + end + + defp flags_from_schema(color?, alpha?, normal?) do + 0 + |> maybe_flag(color?, @flag_color) + |> maybe_flag(alpha?, @flag_alpha) + |> maybe_flag(normal?, @flag_normal) + end + + defp maybe_flag(flags, true, bit), do: Bitwise.bor(flags, bit) + defp maybe_flag(flags, false, _bit), do: flags + + defp schema_flag?(schema, key) when is_list(schema) do + key in schema or Keyword.get(schema, key, false) == true + end + + defp excp_header(flags, count) do + <<@magic::binary, @version::little-unsigned-16, flags::little-unsigned-16, + count::little-unsigned-64>> + end + + defp open_write(path) do + case File.open(path, [:write, :binary, :raw, :read]) do + {:ok, io} -> {:ok, io} + {:error, reason} -> io_error(reason) + end + end + @doc """ Decodes an EXCP version 1 payload into a point cloud. @@ -171,25 +313,40 @@ defmodule ExCodecs.Spatial.Codec.Binary do end @doc """ - Returns an enumerable over points in an EXCP payload. + Returns an enumerable over points in an EXCP payload or file. + + ## Streaming behavior + + * `source: :file` (or `:auto` when the argument looks like a path to a + regular file) reads the 16-byte header, then **one record at a time** + via `IO.binread/2`. Peak memory is O(header + one record), not the + whole cloud. + * `source: :binary` (or `:auto` for an in-memory EXCP payload) still + materializes through `decode/2`, then yields the list. A future Rust + backend may memory-map large binaries instead. + + Prefer `source: :file` for multi‑MB / multi‑GB `.excp` paths. ## Arguments - * `data` (`binary()`) — a complete EXCP payload. - * `opts` (`keyword()`) — reserved and currently ignored. + * `source` (`Path.t() | binary()`) — filesystem path or complete EXCP + payload. Both are binaries, so `:source` controls resolution. + * `opts` (`keyword()`) — `:source` may be `:auto` (default), `:file`, or + `:binary`. ## Returns - An `Enumerable.t()` that yields decoded `%Point{}` values after the complete - cloud has been materialized. If `decode/2` fails, it yields exactly one - `{:error, %ExCodecs.Error{reason: :invalid_data, - codec: :spatial_binary}}` element. + An `Enumerable.t()` yielding `%Point{}` values. Failures are delayed until + enumeration as exactly one `{:error, %ExCodecs.Error{}}`: + + * `reason: :io_error` — the file cannot be opened or read + * `reason: :invalid_data` — bad magic/version, truncated header, or a + truncated record mid-stream (`codec: :spatial_binary`) ## Raises / exceptions - Raises `FunctionClauseError` when `data` is not a binary because this public - function is guarded. `opts` is ignored. Binary validation failures are - delayed as the single error element rather than raised. + Raises `FunctionClauseError` when `source` is not a binary. Unsupported + `:source` values raise `CaseClauseError`. ## Examples @@ -198,24 +355,187 @@ defmodule ExCodecs.Spatial.Codec.Binary do iex> [%Point{x: 0.0, y: 0.0, z: 0.0}] = ...> ExCodecs.Spatial.Codec.Binary.stream_decode(bin) |> Enum.to_list() """ - @spec stream_decode(binary(), keyword()) :: Enumerable.t() - def stream_decode(data, opts \\ []) when is_binary(data) do - case decode(data, opts) do + @spec stream_decode(Path.t() | binary(), keyword()) :: Enumerable.t() + def stream_decode(source, opts \\ []) + + def stream_decode(data, opts) when is_binary(data) do + case resolve_source(data, opts) do + {:ok, :binary, bin} -> + stream_from_binary(bin, opts) + + {:ok, :file, path} -> + Stream.resource( + fn -> open_excp_file(path) end, + &next_excp_item/1, + &close_excp_file/1 + ) + end + end + + defp stream_from_binary(bin, opts) do + case decode(bin, opts) do {:ok, %PointCloud{points: points}} -> Stream.map(points, & &1) {:error, error} -> - Stream.resource( - fn -> {:error, error} end, - fn - {:error, e} -> {[{:error, e}], :done} - :done -> {:halt, :done} - end, - fn _ -> :ok end - ) + error_stream(error) end end + defp resolve_source(bin, opts) do + case Keyword.get(opts, :source, :auto) do + :binary -> + {:ok, :binary, bin} + + :file -> + {:ok, :file, bin} + + :auto -> + if path_like?(bin) and File.regular?(bin) do + {:ok, :file, bin} + else + {:ok, :binary, bin} + end + end + end + + defp path_like?(bin) do + byte_size(bin) < 4096 and not String.starts_with?(bin, @magic) and + (String.contains?(bin, "/") or String.contains?(bin, "\\") or + String.ends_with?(bin, [".excp", ".bin"])) + end + + defp open_excp_file(path) do + case File.open(path, [:read, :binary, :raw]) do + {:ok, io} -> parse_excp_header(io, IO.binread(io, 16)) + {:error, reason} -> io_error(reason) + end + end + + defp parse_excp_header( + io, + <<@magic::binary, version::little-unsigned-16, flags::little-unsigned-16, + count::little-unsigned-64>> + ) do + if version != @version do + File.close(io) + + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Unsupported binary point format version #{version}" + )} + else + {:ok, io, flags, count, record_stride(flags), 0} + end + end + + defp parse_excp_header(io, :eof) do + File.close(io) + + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Invalid ExCodecs binary point cloud" + )} + end + + defp parse_excp_header(io, data) when is_binary(data) do + File.close(io) + + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Invalid ExCodecs binary point cloud" + )} + end + + defp parse_excp_header(io, {:error, reason}) do + File.close(io) + io_error(reason) + end + + defp next_excp_item({:ok, io, _flags, count, _stride, i}) when i >= count do + {:halt, {:done, io}} + end + + defp next_excp_item({:ok, io, flags, count, stride, i}) do + case IO.binread(io, stride) do + data when is_binary(data) and byte_size(data) == stride -> + case decode_point(data, flags) do + {:ok, point, _} -> + {[point], {:ok, io, flags, count, stride, i + 1}} + + {:error, error} -> + {[{:error, error}], {:done, io}} + end + + :eof -> + {[ + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Truncated point record" + )} + ], {:done, io}} + + data when is_binary(data) -> + {[ + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Truncated point record" + )} + ], {:done, io}} + + {:error, reason} -> + {[{:error, io_error_struct(reason)}], {:done, io}} + end + end + + defp next_excp_item({:error, error}), do: {[{:error, error}], :done} + defp next_excp_item({:done, _io}), do: {:halt, :done} + defp next_excp_item(:done), do: {:halt, :done} + + defp close_excp_file({:ok, io, _, _, _, _}), do: File.close(io) + defp close_excp_file({:done, io}), do: File.close(io) + defp close_excp_file(_), do: :ok + + defp record_stride(flags) do + color = + cond do + Bitwise.band(flags, @flag_alpha) != 0 -> 4 + Bitwise.band(flags, @flag_color) != 0 -> 3 + true -> 0 + end + + normal = if Bitwise.band(flags, @flag_normal) != 0, do: 12, else: 0 + 12 + color + normal + end + + defp io_error(reason) do + {:error, io_error_struct(reason)} + end + + defp io_error_struct(reason) do + Error.new(:io_error, + codec: :spatial_binary, + message: "Failed to read EXCP file: #{inspect(reason)}", + details: reason + ) + end + + defp error_stream(error) do + Stream.resource( + fn -> {:error, error} end, + fn + {:error, e} -> {[{:error, e}], :done} + :done -> {:halt, :done} + end, + fn _ -> :ok end + ) + end + defp encode_point(%Point{} = p, flags) do xyz = <> diff --git a/lib/ex_codecs/spatial/codec/gsplat.ex b/lib/ex_codecs/spatial/codec/gsplat.ex index 4d43cee..f885bc0 100644 --- a/lib/ex_codecs/spatial/codec/gsplat.ex +++ b/lib/ex_codecs/spatial/codec/gsplat.ex @@ -95,6 +95,142 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do Error.new(:invalid_data, codec: :gsplat, message: "GSPLAT encode expects a GaussianCloud")} end + @doc """ + Streams `%Gaussian{}` values to a GSPL file using an explicit schema. + + Writes a placeholder header, encodes each Gaussian as it arrives, then seeks + back to patch the final count. Peak memory is O(one Gaussian). + + ## Arguments + + * `enumerable` (`Enumerable.t()`) — `%Gaussian{}` elements. + * `path` (`Path.t()`) — destination file path. + * `opts` (`keyword()`) — requires `:schema`. Use `[]` / `%{}` for no SH + rest coefficients, or `[sh_rest: n]` / `%{sh_rest: n}` for `n` shared + rest floats per Gaussian (shorter lists are zero-padded). + + ## Returns + + * `:ok` + * `{:error, %ExCodecs.Error{reason: :invalid_options}}` when `:schema` is + missing or invalid + * `{:error, %ExCodecs.Error{reason: :invalid_data}}` when an element is not + a `%Gaussian{}` + * `{:error, %ExCodecs.Error{reason: :io_error}}` on file failures + + ## Examples + + iex> alias ExCodecs.Spatial.{Gaussian, Codec.Gsplat} + iex> path = Path.join(System.tmp_dir!(), "gspl_enc_#{System.unique_integer([:positive])}.gspl") + iex> :ok = Gsplat.stream_encode_to_file([Gaussian.new({0, 0, 0})], path, schema: []) + iex> {:ok, <<"GSPL", _::binary>>} = File.read(path) + iex> File.rm!(path) + :ok + """ + @spec stream_encode_to_file(Enumerable.t(), Path.t(), keyword()) :: :ok | {:error, Error.t()} + def stream_encode_to_file(enumerable, path, opts \\ []) do + with {:ok, sh_rest} <- fetch_schema_sh_rest(opts), + {:ok, io} <- open_write(path) do + flags = if sh_rest > 0, do: 1, else: 0 + + try do + :ok = IO.binwrite(io, gspl_header(flags, 0, sh_rest)) + + count = + Enum.reduce(enumerable, 0, fn + %Gaussian{} = g, n -> + :ok = IO.binwrite(io, encode_gaussian(g, sh_rest)) + n + 1 + + other, _n -> + throw({:bad_gaussian, other}) + end) + + {:ok, 0} = :file.position(io, 0) + :ok = IO.binwrite(io, gspl_header(flags, count, sh_rest)) + :ok + catch + {:bad_gaussian, other} -> + {:error, + Error.new(:invalid_data, + codec: :gsplat, + message: "GSPL stream encode expects Gaussian structs, got: #{inspect(other)}" + )} + after + File.close(io) + end + end + end + + defp fetch_schema_sh_rest(opts) do + case Keyword.fetch(opts, :schema) do + :error -> + {:error, + Error.new(:invalid_options, + codec: :gsplat, + message: "GSPL stream_encode_to_file requires schema: (e.g. schema: [sh_rest: 0])" + )} + + {:ok, schema} -> + schema_to_sh_rest(schema) + end + end + + defp schema_to_sh_rest(schema) when is_list(schema) do + sh_rest = Keyword.get(schema, :sh_rest, 0) + + unknown = + schema + |> Enum.reject(fn + {:sh_rest, _} -> true + :sh_rest -> true + _ -> false + end) + + cond do + unknown != [] -> + {:error, + Error.new(:invalid_options, + codec: :gsplat, + message: "Unknown GSPL schema entries: #{inspect(unknown)}" + )} + + not is_integer(sh_rest) or sh_rest < 0 -> + {:error, + Error.new(:invalid_options, + codec: :gsplat, + message: "schema sh_rest must be a non-negative integer" + )} + + true -> + {:ok, sh_rest} + end + end + + defp schema_to_sh_rest(schema) when is_map(schema) do + schema_to_sh_rest(Map.to_list(schema)) + end + + defp schema_to_sh_rest(other) do + {:error, + Error.new(:invalid_options, + codec: :gsplat, + message: "GSPL schema must be a list or map, got: #{inspect(other)}" + )} + end + + defp gspl_header(flags, count, sh_rest) do + <<@magic::binary, @version::little-unsigned-16, flags::little-unsigned-16, + count::little-unsigned-64, sh_rest::little-unsigned-16>> + end + + defp open_write(path) do + case File.open(path, [:write, :binary, :raw, :read]) do + {:ok, io} -> {:ok, io} + {:error, reason} -> io_error(reason) + end + end + @doc """ Decodes a GSPL version 1 payload into a Gaussian cloud. @@ -158,25 +294,39 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do end @doc """ - Returns an enumerable over Gaussians in a GSPL payload. + Returns an enumerable over Gaussians in a GSPL payload or file. + + ## Streaming behavior + + * `source: :file` (or `:auto` when the argument looks like a path to a + regular file) reads the 18-byte header, then **one record at a time** + via `IO.binread/2`. Peak memory is O(header + one record). + * `source: :binary` (or `:auto` for an in-memory GSPL payload) still + materializes through `decode/2`, then yields the list. A future Rust + backend may memory-map large binaries instead. + + Prefer `source: :file` for large `.gspl` paths. ## Arguments - * `data` (`binary()`) — a complete GSPL payload. - * `opts` (`keyword()`) — reserved and currently ignored. + * `source` (`Path.t() | binary()`) — filesystem path or complete GSPL + payload. Both are binaries, so `:source` controls resolution. + * `opts` (`keyword()`) — `:source` may be `:auto` (default), `:file`, or + `:binary`. ## Returns - An `Enumerable.t()` that yields decoded `%Gaussian{}` structs after the - complete cloud has been materialized. If `decode/2` fails, it yields exactly - one `{:error, %ExCodecs.Error{reason: :invalid_data, codec: :gsplat}}` - element. + An `Enumerable.t()` yielding `%Gaussian{}` values. Failures are delayed until + enumeration as exactly one `{:error, %ExCodecs.Error{}}`: + + * `reason: :io_error` — the file cannot be opened or read + * `reason: :invalid_data` — bad magic/version, truncated header, or a + truncated record mid-stream (`codec: :gsplat`) ## Raises / exceptions - Raises `FunctionClauseError` when `data` is not a binary because this public - function is guarded. `opts` is ignored. Binary validation failures are - delayed as the single error element rather than raised. + Raises `FunctionClauseError` when `source` is not a binary. Unsupported + `:source` values raise `CaseClauseError`. ## Examples @@ -185,24 +335,157 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do iex> [%Gaussian{opacity: 1.0}] = ...> ExCodecs.Spatial.Codec.Gsplat.stream_decode(bin) |> Enum.to_list() """ - @spec stream_decode(binary(), keyword()) :: Enumerable.t() - def stream_decode(data, opts \\ []) when is_binary(data) do - case decode(data, opts) do + @spec stream_decode(Path.t() | binary(), keyword()) :: Enumerable.t() + def stream_decode(source, opts \\ []) + + def stream_decode(data, opts) when is_binary(data) do + case resolve_source(data, opts) do + {:ok, :binary, bin} -> + stream_from_binary(bin, opts) + + {:ok, :file, path} -> + Stream.resource( + fn -> open_gspl_file(path) end, + &next_gspl_item/1, + &close_gspl_file/1 + ) + end + end + + defp stream_from_binary(bin, opts) do + case decode(bin, opts) do {:ok, %GaussianCloud{gaussians: gs}} -> Stream.map(gs, & &1) {:error, error} -> - Stream.resource( - fn -> {:error, error} end, - fn - {:error, e} -> {[{:error, e}], :done} - :done -> {:halt, :done} - end, - fn _ -> :ok end - ) + error_stream(error) + end + end + + defp resolve_source(bin, opts) do + case Keyword.get(opts, :source, :auto) do + :binary -> + {:ok, :binary, bin} + + :file -> + {:ok, :file, bin} + + :auto -> + if path_like?(bin) and File.regular?(bin) do + {:ok, :file, bin} + else + {:ok, :binary, bin} + end + end + end + + defp path_like?(bin) do + byte_size(bin) < 4096 and not String.starts_with?(bin, @magic) and + (String.contains?(bin, "/") or String.contains?(bin, "\\") or + String.ends_with?(bin, [".gspl", ".bin"])) + end + + defp open_gspl_file(path) do + case File.open(path, [:read, :binary, :raw]) do + {:ok, io} -> parse_gspl_header(io, IO.binread(io, 18)) + {:error, reason} -> io_error(reason) + end + end + + defp parse_gspl_header( + io, + <<@magic::binary, version::little-unsigned-16, _flags::little-unsigned-16, + count::little-unsigned-64, sh_rest::little-unsigned-16>> + ) do + if version != @version do + File.close(io) + + {:error, + Error.new(:invalid_data, + codec: :gsplat, + message: "Unsupported GSPLAT version #{version}" + )} + else + {:ok, io, sh_rest, count, record_stride(sh_rest), 0} end end + defp parse_gspl_header(io, :eof) do + File.close(io) + {:error, Error.new(:invalid_data, codec: :gsplat, message: "Invalid GSPLAT binary")} + end + + defp parse_gspl_header(io, data) when is_binary(data) do + File.close(io) + {:error, Error.new(:invalid_data, codec: :gsplat, message: "Invalid GSPLAT binary")} + end + + defp parse_gspl_header(io, {:error, reason}) do + File.close(io) + io_error(reason) + end + + defp next_gspl_item({:ok, io, _sh_rest, count, _stride, i}) when i >= count do + {:halt, {:done, io}} + end + + defp next_gspl_item({:ok, io, sh_rest, count, stride, i}) do + case IO.binread(io, stride) do + data when is_binary(data) and byte_size(data) == stride -> + case decode_gaussian(data, sh_rest) do + {:ok, gaussian, _} -> + {[gaussian], {:ok, io, sh_rest, count, stride, i + 1}} + + {:error, error} -> + {[{:error, error}], {:done, io}} + end + + :eof -> + {[ + {:error, Error.new(:invalid_data, codec: :gsplat, message: "Truncated Gaussian record")} + ], {:done, io}} + + data when is_binary(data) -> + {[ + {:error, Error.new(:invalid_data, codec: :gsplat, message: "Truncated Gaussian record")} + ], {:done, io}} + + {:error, reason} -> + {[{:error, io_error_struct(reason)}], {:done, io}} + end + end + + defp next_gspl_item({:error, error}), do: {[{:error, error}], :done} + defp next_gspl_item({:done, _io}), do: {:halt, :done} + defp next_gspl_item(:done), do: {:halt, :done} + + defp close_gspl_file({:ok, io, _, _, _, _}), do: File.close(io) + defp close_gspl_file({:done, io}), do: File.close(io) + defp close_gspl_file(_), do: :ok + + defp record_stride(sh_rest), do: 56 + sh_rest * 4 + + defp io_error(reason), do: {:error, io_error_struct(reason)} + + defp io_error_struct(reason) do + Error.new(:io_error, + codec: :gsplat, + message: "Failed to read GSPL file: #{inspect(reason)}", + details: reason + ) + end + + defp error_stream(error) do + Stream.resource( + fn -> {:error, error} end, + fn + {:error, e} -> {[{:error, e}], :done} + :done -> {:halt, :done} + end, + fn _ -> :ok end + ) + end + defp max_sh_rest(gaussians) do gaussians |> Enum.map(fn diff --git a/lib/ex_codecs/spatial/codec/ply.ex b/lib/ex_codecs/spatial/codec/ply.ex index 18c07ab..0872b74 100644 --- a/lib/ex_codecs/spatial/codec/ply.ex +++ b/lib/ex_codecs/spatial/codec/ply.ex @@ -195,8 +195,13 @@ defmodule ExCodecs.Spatial.Codec.PLY do @doc """ Returns an enumerable over vertices from a PLY path or binary. - Both modes currently read and decode the entire payload before yielding - elements; this is not incremental PLY parsing. + ## Streaming behavior + + * `source: :file` (or `:auto` path detection): scans the header until + `end_header`, then yields one vertex at a time — binary formats via + fixed-stride `IO.binread/2`, ASCII via line reads. Peak memory is + O(header + one vertex / read chunk), not the whole cloud. + * `source: :binary`: still materializes through `decode/2`, then yields. ## Arguments @@ -237,9 +242,9 @@ defmodule ExCodecs.Spatial.Codec.PLY do {:ok, :file, path} -> Stream.resource( - fn -> open_ply_file(path, opts) end, - &next_ply_item/1, - &close_ply_file/1 + fn -> open_ply_stream(path, opts) end, + &next_ply_stream_item/1, + &close_ply_stream/1 ) end end @@ -773,38 +778,37 @@ defmodule ExCodecs.Spatial.Codec.PLY do defp decode_gaussians(body, parsed) do with {:ok, points} <- decode_vertices(body, parsed) do - gaussians = - Enum.map(points, fn %Point{} = p -> - # Point.new/4 normalizes attribute keys to strings. - attrs = p.attributes - - Gaussian.new({p.x, p.y, p.z}, - color: { - Map.get(attrs, "f_dc_0", 0.5), - Map.get(attrs, "f_dc_1", 0.5), - Map.get(attrs, "f_dc_2", 0.5) - }, - opacity: Map.get(attrs, "opacity", 1.0), - scale: { - Map.get(attrs, "scale_0", 1.0), - Map.get(attrs, "scale_1", 1.0), - Map.get(attrs, "scale_2", 1.0) - }, - rotation: { - Map.get(attrs, "rot_0", 1.0), - Map.get(attrs, "rot_1", 0.0), - Map.get(attrs, "rot_2", 0.0), - Map.get(attrs, "rot_3", 0.0) - }, - sh: extract_sh(attrs), - metadata: attrs - ) - end) - - {:ok, gaussians} + {:ok, Enum.map(points, &point_to_gaussian/1)} end end + defp point_to_gaussian(%Point{} = p) do + # Point.new/4 normalizes attribute keys to strings. + attrs = p.attributes + + Gaussian.new({p.x, p.y, p.z}, + color: { + Map.get(attrs, "f_dc_0", 0.5), + Map.get(attrs, "f_dc_1", 0.5), + Map.get(attrs, "f_dc_2", 0.5) + }, + opacity: Map.get(attrs, "opacity", 1.0), + scale: { + Map.get(attrs, "scale_0", 1.0), + Map.get(attrs, "scale_1", 1.0), + Map.get(attrs, "scale_2", 1.0) + }, + rotation: { + Map.get(attrs, "rot_0", 1.0), + Map.get(attrs, "rot_1", 0.0), + Map.get(attrs, "rot_2", 0.0), + Map.get(attrs, "rot_3", 0.0) + }, + sh: extract_sh(attrs), + metadata: attrs + ) + end + defp extract_sh(attrs) do rest_keys = attrs @@ -923,6 +927,9 @@ defmodule ExCodecs.Spatial.Codec.PLY do # --- Streaming helpers ---------------------------------------------------- + @header_scan_chunk 4096 + @header_max_bytes 1_048_576 + defp decode_to_list(data, opts) when is_binary(data) do case decode(data, opts) do {:ok, %PointCloud{points: points}} -> {:ok, points} @@ -931,30 +938,224 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end - defp open_ply_file(path, opts) do - case File.read(path) do - {:ok, data} -> - case decode_to_list(data, opts) do - {:ok, items} -> {:ok, items} - {:error, error} -> {:error, error} + defp open_ply_stream(path, opts) do + case File.open(path, [:read, :binary, :raw]) do + {:ok, io} -> + case scan_ply_header(io, "") do + {:ok, parsed, leftover} -> + as = resolve_as(Keyword.get(opts, :as, :auto), parsed.properties) + {:ok, init_ply_stream_state(io, leftover, parsed, as)} + + {:error, error} -> + File.close(io) + {:error, error} + end + + {:error, reason} -> + ply_io_error(reason) + end + end + + defp init_ply_stream_state(io, leftover, %{format: :ascii} = parsed, as) do + {:ascii, io, leftover, parsed, as, 0} + end + + defp init_ply_stream_state(io, leftover, %{format: endian} = parsed, as) + when endian in [:binary_le, :binary_be] do + stride = Enum.reduce(parsed.properties, 0, fn p, acc -> acc + type_size(p.type) end) + {:binary, io, leftover, endian, parsed, as, stride, 0} + end + + defp scan_ply_header(io, acc) do + case IO.binread(io, @header_scan_chunk) do + data when is_binary(data) and byte_size(data) > 0 -> + buf = acc <> data + match_or_continue_header(io, buf) + + :eof -> + header_eof_error(acc) + + {:error, reason} -> + ply_io_error(reason) + + _other -> + header_eof_error(acc) + end + end + + defp match_or_continue_header(io, buf) do + case :binary.match(buf, "end_header") do + {pos, len} -> + header = binary_part(buf, 0, pos + len) + rest = binary_part(buf, pos + len, byte_size(buf) - pos - len) + + with {:ok, parsed} <- parse_header(header) do + {:ok, parsed, strip_leading_newlines(rest)} + end + + :nomatch when byte_size(buf) > @header_max_bytes -> + {:error, + Error.new(:invalid_data, + codec: :ply, + message: "PLY header exceeds #{@header_max_bytes} bytes without end_header" + )} + + :nomatch -> + scan_ply_header(io, buf) + end + end + + defp header_eof_error("") do + {:error, Error.new(:invalid_data, codec: :ply, message: "PLY header missing end_header")} + end + + defp header_eof_error(acc) do + case :binary.match(acc, "end_header") do + {pos, len} -> + header = binary_part(acc, 0, pos + len) + rest = binary_part(acc, pos + len, byte_size(acc) - pos - len) + + with {:ok, parsed} <- parse_header(header) do + {:ok, parsed, strip_leading_newlines(rest)} + end + + :nomatch -> + {:error, Error.new(:invalid_data, codec: :ply, message: "PLY header missing end_header")} + end + end + + defp next_ply_stream_item({:ok, {:binary, io, _buf, _endian, parsed, _as, _stride, i}}) + when i >= parsed.count do + {:halt, {:done, io}} + end + + defp next_ply_stream_item({:ok, {:binary, io, buf, endian, parsed, as, stride, i}}) do + case take_exact_bytes(io, buf, stride) do + {:ok, record, rest} -> + {values, _} = unpack_row(record, parsed.properties, endian) + item = emit_vertex(values_to_point(values, parsed.properties), as) + {[item], {:ok, {:binary, io, rest, endian, parsed, as, stride, i + 1}}} + + {:error, error} -> + {[{:error, error}], {:done, io}} + end + end + + defp next_ply_stream_item({:ok, {:ascii, io, _buf, parsed, _as, i}}) when i >= parsed.count do + {:halt, {:done, io}} + end + + defp next_ply_stream_item({:ok, {:ascii, io, buf, parsed, as, i}}) do + case take_ascii_vertex_line(io, buf) do + {:ok, line, rest} -> + values = line |> String.split() |> Enum.map(&parse_ascii_number/1) + item = emit_vertex(values_to_point(values, parsed.properties), as) + {[item], {:ok, {:ascii, io, rest, parsed, as, i + 1}}} + + {:error, error} -> + {[{:error, error}], {:done, io}} + end + end + + defp next_ply_stream_item({:error, error}), do: {[{:error, error}], :done} + defp next_ply_stream_item({:done, _io}), do: {:halt, :done} + defp next_ply_stream_item(:done), do: {:halt, :done} + + defp close_ply_stream({:ok, {:binary, io, _, _, _, _, _, _}}), do: File.close(io) + defp close_ply_stream({:ok, {:ascii, io, _, _, _, _}}), do: File.close(io) + defp close_ply_stream({:done, io}), do: File.close(io) + defp close_ply_stream(_), do: :ok + + defp emit_vertex(point, :point_cloud), do: point + defp emit_vertex(point, :gaussian_cloud), do: point_to_gaussian(point) + + defp take_exact_bytes(_io, buf, n) when byte_size(buf) >= n do + <> = buf + {:ok, chunk, rest} + end + + defp take_exact_bytes(io, buf, n) do + case IO.binread(io, n - byte_size(buf)) do + data when is_binary(data) and byte_size(data) > 0 -> + take_exact_bytes(io, buf <> data, n) + + :eof -> + {:error, + Error.new(:invalid_data, codec: :ply, message: "Binary PLY body too short")} + + data when is_binary(data) -> + {:error, + Error.new(:invalid_data, codec: :ply, message: "Binary PLY body too short")} + + {:error, reason} -> + ply_io_error(reason) + end + end + + defp take_ascii_vertex_line(io, buf) do + case split_first_line(buf) do + {:ok, line, rest} -> + if String.trim(line) == "" do + take_ascii_vertex_line(io, rest) + else + {:ok, String.trim(line), rest} + end + + :incomplete -> + read_more_ascii_line(io, buf) + end + end + + defp read_more_ascii_line(io, buf) do + case IO.binread(io, @header_scan_chunk) do + data when is_binary(data) and byte_size(data) > 0 -> + take_ascii_vertex_line(io, buf <> data) + + :eof -> + trimmed = String.trim(buf) + + if trimmed == "" do + {:error, + Error.new(:invalid_data, + codec: :ply, + message: "Expected more ASCII vertices" + )} + else + {:ok, trimmed, ""} end {:error, reason} -> + ply_io_error(reason) + + _other -> {:error, - Error.new(:io_error, + Error.new(:invalid_data, codec: :ply, - message: "Failed to read PLY file: #{inspect(reason)}", - details: reason + message: "Expected more ASCII vertices" )} end end - defp next_ply_item({:ok, [item | rest]}), do: {[item], {:ok, rest}} - defp next_ply_item({:ok, []}), do: {:halt, :done} - defp next_ply_item({:error, error}), do: {[{:error, error}], :done} - defp next_ply_item(:done), do: {:halt, :done} + defp split_first_line(buf) do + case :binary.match(buf, "\n") do + {pos, 1} -> + line = binary_part(buf, 0, pos) |> String.trim_trailing("\r") + rest = binary_part(buf, pos + 1, byte_size(buf) - pos - 1) + {:ok, line, rest} - defp close_ply_file(_), do: :ok + :nomatch -> + :incomplete + end + end + + defp ply_io_error(reason) do + {:error, + Error.new(:io_error, + codec: :ply, + message: "Failed to read PLY file: #{inspect(reason)}", + details: reason + )} + end defp error_stream({:error, error}), do: {[{:error, error}], :done} defp error_stream(:done), do: {:halt, :done} diff --git a/lib/ex_codecs/spatial/stream.ex b/lib/ex_codecs/spatial/stream.ex index ac35dec..7a071ba 100644 --- a/lib/ex_codecs/spatial/stream.ex +++ b/lib/ex_codecs/spatial/stream.ex @@ -2,9 +2,17 @@ defmodule ExCodecs.Spatial.Stream do @moduledoc """ Stream helpers for spatial formats. - **Important:** despite the name, most paths **materialize** the full source - (or full enumerable) then yield items. True incremental I/O for multi-GB - files is not implemented yet. + ## What actually streams today + + * **EXCP** (`:spatial_binary`), **GSPL** (`:gsplat`), and **PLY** with + `source: :file` (or `:auto` path detection) read the header, then one + record/vertex at a time from disk — O(header + one record) memory. + Binary PLY uses a fixed stride; ASCII PLY reads lines. + * In-memory binaries still **materialize** through `decode/2`, then yield. + A future Rust backend may memory-map large binaries. + * **`encode_to_file/3`** with an explicit `:schema` streams EXCP/GSPL to + disk (placeholder header, then seek-back count patch). Without `:schema`, + encoding still collects in memory first. Prefer `source: :file` when the argument is a filesystem path, or `source: :binary` when it is an encoded payload. With `:auto` (default), a @@ -22,8 +30,8 @@ defmodule ExCodecs.Spatial.Stream do @doc """ Returns an enumerable over points or Gaussians decoded from a path or binary. - The complete source and decoded cloud are currently materialized before - elements are emitted. + EXCP/GSPL/PLY file sources stream incrementally; in-memory binaries still + materialize before emitting elements. ## Arguments @@ -42,7 +50,7 @@ defmodule ExCodecs.Spatial.Stream do * `reason: :invalid_options` — `:format` is absent. * `reason: :unsupported_codec` — `:format` is unknown. - * `reason: :io_error` — a selected file cannot be read. + * `reason: :io_error` — a selected file cannot be opened or read. * `reason: :invalid_data` — the payload is malformed, unsupported, or truncated. @@ -184,15 +192,22 @@ defmodule ExCodecs.Spatial.Stream do @doc """ Encodes spatial data and writes the binary to `path`. + When `:schema` is present and `:format` is `:spatial_binary` or `:gsplat`, + points/Gaussians are written incrementally via + `Binary.stream_encode_to_file/3` / `Gsplat.stream_encode_to_file/3` + (placeholder header + seek-back count). Otherwise the payload is encoded in + memory and written with `File.write/2`. + ## Arguments * `data` (`PointCloud.t() | GaussianCloud.t() | Enumerable.t()`) — a cloud, or an enumerable of `%Point{}`/`%Gaussian{}` values. - * `path` (`Path.t()`) — destination accepted by `File.write/2`; parent - directories must already exist. + * `path` (`Path.t()`) — destination path; parent directories must already + exist. * `opts` (`keyword()`) — the options for `ExCodecs.Spatial.encode/2`, with `:format` required for enumerable input and defaulting to `:ply` for an - already-built cloud. + already-built cloud. For streaming EXCP/GSPL writes, pass `:schema` + (see codec docs). ## Returns @@ -200,15 +215,14 @@ defmodule ExCodecs.Spatial.Stream do * any `:invalid_options`, `:unsupported_codec`, or `:invalid_data` error returned by encoding. * `{:error, %ExCodecs.Error{reason: :io_error, details: reason}}` when - `File.write/2` returns a file/POSIX error such as `:enoent`, `:eacces`, - or `:enospc`. + a file/POSIX error occurs (e.g. `:enoent`, `:eacces`, `:enospc`). ## Raises / exceptions - Normal `File.write/2` path/POSIX failures are returned. Invalid path terms or - non-path binaries can raise `FunctionClauseError` or `ArgumentError` in the - path/file APIs. Encoding also has the enumerable, keyword, and - malformed-struct exceptions documented by `encode/2`. + Normal path/POSIX failures are returned. Invalid path terms or non-path + binaries can raise `FunctionClauseError` or `ArgumentError` in the path/file + APIs. Encoding also has the enumerable, keyword, and malformed-struct + exceptions documented by `encode/2`. ## Examples @@ -222,6 +236,50 @@ defmodule ExCodecs.Spatial.Stream do @spec encode_to_file(Enumerable.t() | PointCloud.t() | GaussianCloud.t(), Path.t(), keyword()) :: :ok | {:error, Error.t()} def encode_to_file(data, path, opts \\ []) do + case stream_encode_to_file_if_schema(data, path, opts) do + :not_streaming -> + write_encoded_payload(data, path, opts) + + result -> + result + end + end + + defp stream_encode_to_file_if_schema(data, path, opts) do + if Keyword.has_key?(opts, :schema) do + case Keyword.get(opts, :format) do + :spatial_binary -> + Binary.stream_encode_to_file(enumerable_points(data), path, opts) + + :gsplat -> + Gsplat.stream_encode_to_file(enumerable_gaussians(data), path, opts) + + other when other in [:ply, nil] -> + {:error, + Error.new(:invalid_options, + message: + "encode_to_file schema streaming requires format: :spatial_binary or :gsplat" + )} + + other -> + {:error, + Error.new(:unsupported_codec, + codec: other, + message: "Unsupported spatial stream format: #{inspect(other)}" + )} + end + else + :not_streaming + end + end + + defp enumerable_points(%PointCloud{points: points}), do: points + defp enumerable_points(enumerable), do: enumerable + + defp enumerable_gaussians(%GaussianCloud{gaussians: gs}), do: gs + defp enumerable_gaussians(enumerable), do: enumerable + + defp write_encoded_payload(data, path, opts) do result = case data do %PointCloud{} -> ExCodecs.Spatial.encode(data, opts) @@ -256,56 +314,9 @@ defmodule ExCodecs.Spatial.Stream do {:error, Error.new(:unsupported_codec, codec: other)} end - defp stream_binary(source, opts) do - case resolve_source(source, opts) do - {:ok, bin} -> Binary.stream_decode(bin, opts) - {:error, error} -> error_stream(error) - end - end - - defp stream_gsplat(source, opts) do - case resolve_source(source, opts) do - {:ok, bin} -> Gsplat.stream_decode(bin, opts) - {:error, error} -> error_stream(error) - end - end - - defp resolve_source(bin, opts) when is_binary(bin) do - case Keyword.get(opts, :source, :auto) do - :binary -> - {:ok, bin} - - :file -> - read_file(bin) - - :auto -> - if path_like?(bin) and File.regular?(bin) do - read_file(bin) - else - {:ok, bin} - end - end - end - - defp path_like?(bin) do - byte_size(bin) < 4096 and - not String.starts_with?(bin, "EXCP") and - not String.starts_with?(bin, "GSPL") and - not String.starts_with?(bin, "ply") and - (String.contains?(bin, "/") or String.contains?(bin, "\\") or - String.ends_with?(bin, [".excp", ".gspl", ".bin", ".ply"])) - end - - defp read_file(path) do - case File.read(path) do - {:ok, data} -> - {:ok, data} - - {:error, reason} -> - {:error, - Error.new(:io_error, message: "Failed to read file: #{inspect(reason)}", details: reason)} - end - end + # EXCP/GSPL resolve :source and stream from disk themselves (no full File.read). + defp stream_binary(source, opts), do: Binary.stream_decode(source, opts) + defp stream_gsplat(source, opts), do: Gsplat.stream_decode(source, opts) defp error_stream(error) do Stream.resource( diff --git a/test/ex_codecs/spatial/codec/binary_test.exs b/test/ex_codecs/spatial/codec/binary_test.exs index aa9788c..fae9929 100644 --- a/test/ex_codecs/spatial/codec/binary_test.exs +++ b/test/ex_codecs/spatial/codec/binary_test.exs @@ -30,4 +30,81 @@ defmodule ExCodecs.Spatial.Codec.BinaryTest do assert {:error, %{reason: :invalid_data}} = Binary.decode(<<"XXXX", 0, 1, 0, 0, 0, 0, 0, 0, 0, 0>>) end + + test "stream_decode from file yields points incrementally" do + cloud = + PointCloud.new([ + Point.new(1.0, 2.0, 3.0, color: {10, 20, 30}), + Point.new(4.0, 5.0, 6.0, color: {40, 50, 60}) + ]) + + assert {:ok, bin} = Binary.encode(cloud) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_excp_stream_#{System.unique_integer([:positive])}.excp" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + points = Binary.stream_decode(path, source: :file) |> Enum.to_list() + assert length(points) == 2 + assert hd(points).color == {10, 20, 30} + + # :auto path detection + assert [%Point{}, %Point{}] = + Binary.stream_decode(path) |> Enum.to_list() + end + + test "stream_decode from truncated file yields invalid_data" do + cloud = PointCloud.new([Point.new(1.0, 2.0, 3.0), Point.new(4.0, 5.0, 6.0)]) + assert {:ok, bin} = Binary.encode(cloud) + # Header (16) + one incomplete xyz record + truncated = binary_part(bin, 0, 20) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_excp_trunc_#{System.unique_integer([:positive])}.excp" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, truncated) + + assert [{:error, %{reason: :invalid_data, codec: :spatial_binary}}] = + Binary.stream_decode(path, source: :file) |> Enum.to_list() + end + + test "stream_decode missing file yields io_error" do + assert [{:error, %{reason: :io_error, codec: :spatial_binary}}] = + Binary.stream_decode("/no/such/ex_codecs.excp", source: :file) |> Enum.to_list() + end + + test "stream_encode_to_file with schema round-trips" do + points = [ + Point.new(1.0, 2.0, 3.0, color: {10, 20, 30}), + Point.new(4.0, 5.0, 6.0, color: {40, 50, 60}) + ] + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_excp_enc_#{System.unique_integer([:positive])}.excp" + ) + + on_exit(fn -> File.rm(path) end) + + assert :ok = Binary.stream_encode_to_file(points, path, schema: [:color]) + assert {:ok, <<"EXCP", _::binary>> = bin} = File.read(path) + assert {:ok, decoded} = Binary.decode(bin) + assert length(decoded.points) == 2 + assert hd(decoded.points).color == {10, 20, 30} + end + + test "stream_encode_to_file requires schema" do + assert {:error, %{reason: :invalid_options}} = + Binary.stream_encode_to_file([Point.new(0, 0, 0)], "x.excp") + end end diff --git a/test/ex_codecs/spatial/codec/gsplat_test.exs b/test/ex_codecs/spatial/codec/gsplat_test.exs index 15e38d9..efc9fa0 100644 --- a/test/ex_codecs/spatial/codec/gsplat_test.exs +++ b/test/ex_codecs/spatial/codec/gsplat_test.exs @@ -38,4 +38,75 @@ defmodule ExCodecs.Spatial.Codec.GsplatTest do assert is_list(g.sh) assert length(List.flatten(tl(g.sh))) == 6 end + + test "stream_decode from file yields Gaussians incrementally" do + cloud = + GaussianCloud.new([ + Gaussian.new({1.0, 0.0, 0.0}, opacity: 0.5), + Gaussian.new({0.0, 1.0, 0.0}, + color: {0.2, 0.3, 0.4}, + sh: [[0.2, 0.3, 0.4], [0.1, 0.1, 0.1]] + ) + ]) + + assert {:ok, bin} = Gsplat.encode(cloud) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_gspl_stream_#{System.unique_integer([:positive])}.gspl" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + gs = Gsplat.stream_decode(path, source: :file) |> Enum.to_list() + assert length(gs) == 2 + assert_in_delta hd(gs).opacity, 0.5, 0.0001 + assert is_list(List.last(gs).sh) + end + + test "stream_decode from truncated file yields invalid_data" do + cloud = GaussianCloud.new([Gaussian.new({1.0, 2.0, 3.0}), Gaussian.new({4.0, 5.0, 6.0})]) + assert {:ok, bin} = Gsplat.encode(cloud) + truncated = binary_part(bin, 0, 30) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_gspl_trunc_#{System.unique_integer([:positive])}.gspl" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, truncated) + + assert [{:error, %{reason: :invalid_data, codec: :gsplat}}] = + Gsplat.stream_decode(path, source: :file) |> Enum.to_list() + end + + test "stream_encode_to_file with schema round-trips" do + gs = [ + Gaussian.new({1.0, 2.0, 3.0}, opacity: 0.5), + Gaussian.new({0.0, 1.0, 0.0}, color: {0.2, 0.3, 0.4}) + ] + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_gspl_enc_#{System.unique_integer([:positive])}.gspl" + ) + + on_exit(fn -> File.rm(path) end) + + assert :ok = Gsplat.stream_encode_to_file(gs, path, schema: []) + assert {:ok, <<"GSPL", _::binary>> = bin} = File.read(path) + assert {:ok, decoded} = Gsplat.decode(bin) + assert length(decoded.gaussians) == 2 + assert_in_delta hd(decoded.gaussians).opacity, 0.5, 0.0001 + end + + test "stream_encode_to_file requires schema" do + assert {:error, %{reason: :invalid_options}} = + Gsplat.stream_encode_to_file([Gaussian.new({0, 0, 0})], "x.gspl") + end end diff --git a/test/ex_codecs/spatial/codec/ply_test.exs b/test/ex_codecs/spatial/codec/ply_test.exs index 88c2b30..70633cc 100644 --- a/test/ex_codecs/spatial/codec/ply_test.exs +++ b/test/ex_codecs/spatial/codec/ply_test.exs @@ -92,4 +92,89 @@ defmodule ExCodecs.Spatial.Codec.PLYTest do test "rejects non-PLY data" do assert {:error, %{reason: :invalid_data}} = PLY.decode("not a ply file") end + + describe "stream_decode from file" do + test "streams ASCII vertices without materializing the whole body" do + cloud = + PointCloud.new([ + Point.new(1.0, 2.0, 3.0, color: {255, 0, 0}), + Point.new(4.0, 5.0, 6.0, color: {0, 255, 0}) + ]) + + assert {:ok, bin} = PLY.encode(cloud, format: :ascii) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_ply_ascii_stream_#{System.unique_integer([:positive])}.ply" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + points = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert length(points) == 2 + assert hd(points).color == {255, 0, 0} + assert_in_delta List.last(points).x, 4.0, 0.0001 + end + + test "streams binary little-endian vertices" do + cloud = + PointCloud.new([ + Point.new(1.25, 2.5, 3.75, color: {1, 2, 3}), + Point.new(-1.0, 0.0, 9.0, color: {4, 5, 6}) + ]) + + assert {:ok, bin} = PLY.encode(cloud, ply_format: :binary_le) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_ply_bin_stream_#{System.unique_integer([:positive])}.ply" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + points = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert length(points) == 2 + assert_in_delta hd(points).x, 1.25, 0.0001 + assert hd(points).color == {1, 2, 3} + end + + test "streams Gaussian ASCII PLY as Gaussians" do + cloud = GaussianCloud.new([Gaussian.new({1.0, 2.0, 3.0}, opacity: 0.7)]) + assert {:ok, bin} = PLY.encode(cloud, format: :ascii) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_ply_gauss_stream_#{System.unique_integer([:positive])}.ply" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + assert [%Gaussian{} = g] = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert_in_delta g.opacity, 0.7, 0.0001 + end + + test "truncated binary body yields invalid_data" do + cloud = PointCloud.new([Point.new(1.0, 2.0, 3.0), Point.new(4.0, 5.0, 6.0)]) + assert {:ok, bin} = PLY.encode(cloud, ply_format: :binary_le) + truncated = binary_part(bin, 0, byte_size(bin) - 4) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_ply_trunc_#{System.unique_integer([:positive])}.ply" + ) + + on_exit(fn -> File.rm(path) end) + File.write!(path, truncated) + + items = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert Enum.any?(items, &match?({:error, %{reason: :invalid_data, codec: :ply}}, &1)) + end + end end diff --git a/test/ex_codecs/spatial/stream_test.exs b/test/ex_codecs/spatial/stream_test.exs index 879479f..d2a44aa 100644 --- a/test/ex_codecs/spatial/stream_test.exs +++ b/test/ex_codecs/spatial/stream_test.exs @@ -66,4 +66,37 @@ defmodule ExCodecs.Spatial.StreamTest do {:ok, bin} = Spatial.encode(cloud, format: :ply) assert [%Point{}] = ExCodecs.stream_decode(bin, format: :ply) |> Enum.to_list() end + + test "encode_to_file with schema streams EXCP" do + alias ExCodecs.Spatial.Stream, as: SpatialStream + + points = [Point.new(1.0, 2.0, 3.0, color: {1, 2, 3})] + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_stream_enc_#{System.unique_integer([:positive])}.excp" + ) + + on_exit(fn -> File.rm(path) end) + + assert :ok = + SpatialStream.encode_to_file(points, path, + format: :spatial_binary, + schema: [:color] + ) + + assert [%Point{color: {1, 2, 3}}] = + Spatial.stream_decode(path, format: :spatial_binary, source: :file) |> Enum.to_list() + end + + test "encode_to_file schema requires EXCP or GSPL format" do + alias ExCodecs.Spatial.Stream, as: SpatialStream + + assert {:error, %{reason: :invalid_options}} = + SpatialStream.encode_to_file([Point.new(0, 0, 0)], "x.ply", + format: :ply, + schema: [] + ) + end end From 7ac0016e13138ec31373cc42429f98e37d93a8b8 Mon Sep 17 00:00:00 2001 From: thanos Date: Fri, 17 Jul 2026 17:50:08 -0400 Subject: [PATCH 2/6] Split stream_encode_to_file_if_schema/3 into format-specific clauses plus enumerable_points/1 / enumerable_gaussians/1. Credo is clean and the stream tests pass. --- lib/ex_codecs/spatial/codec/binary.ex | 39 +- lib/ex_codecs/spatial/codec/gsplat.ex | 20 +- lib/ex_codecs/spatial/codec/ply.ex | 50 +- lib/ex_codecs/spatial/stream.ex | 45 +- .../spatial/stream_coverage_test.exs | 461 ++++++++++++++++++ 5 files changed, 534 insertions(+), 81 deletions(-) create mode 100644 test/ex_codecs/spatial/stream_coverage_test.exs diff --git a/lib/ex_codecs/spatial/codec/binary.ex b/lib/ex_codecs/spatial/codec/binary.ex index b7c3457..d8744cd 100644 --- a/lib/ex_codecs/spatial/codec/binary.ex +++ b/lib/ex_codecs/spatial/codec/binary.ex @@ -129,6 +129,13 @@ defmodule ExCodecs.Spatial.Codec.Binary do a `%Point{}` * `{:error, %ExCodecs.Error{reason: :io_error}}` on file failures + ## Raises / exceptions + + Schema and non-`%Point{}` element failures are returned. Invalid path terms or + non-keyword `opts` can raise `FunctionClauseError` or `ArgumentError` in the + path/file/keyword APIs. Malformed point field shapes can raise during + bitstring construction (same class of exceptions as `encode/2`). + ## Examples iex> alias ExCodecs.Spatial.{Point, Codec.Binary} @@ -190,13 +197,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do color? = alpha? or schema_flag?(schema, :color) normal? = schema_flag?(schema, :normal) - unknown = - schema - |> Enum.reject(fn - a when is_atom(a) -> a in [:color, :alpha, :normal] - {k, _} when is_atom(k) -> k in [:color, :alpha, :normal] - _ -> false - end) + unknown = Enum.reject(schema, &known_excp_schema_entry?/1) if unknown != [] do {:error, @@ -221,6 +222,10 @@ defmodule ExCodecs.Spatial.Codec.Binary do )} end + defp known_excp_schema_entry?(entry) when entry in [:color, :alpha, :normal], do: true + defp known_excp_schema_entry?({key, _}) when key in [:color, :alpha, :normal], do: true + defp known_excp_schema_entry?(_), do: false + defp flags_from_schema(color?, alpha?, normal?) do 0 |> maybe_flag(color?, @flag_color) @@ -564,13 +569,21 @@ defmodule ExCodecs.Spatial.Codec.Binary do xyz <> color <> normal end - defp normalize_rgb(nil), do: {0, 0, 0} - defp normalize_rgb({r, g, b}), do: {trunc(r), trunc(g), trunc(b)} - defp normalize_rgb({r, g, b, _a}), do: {trunc(r), trunc(g), trunc(b)} + defp normalize_rgb(color) do + case color do + nil -> {0, 0, 0} + {r, g, b} -> {trunc(r), trunc(g), trunc(b)} + {r, g, b, _a} -> {trunc(r), trunc(g), trunc(b)} + end + end - defp normalize_rgba(nil), do: {0, 0, 0, 255} - defp normalize_rgba({r, g, b}), do: {trunc(r), trunc(g), trunc(b), 255} - defp normalize_rgba({r, g, b, a}), do: {trunc(r), trunc(g), trunc(b), trunc(a)} + defp normalize_rgba(color) do + case color do + nil -> {0, 0, 0, 255} + {r, g, b} -> {trunc(r), trunc(g), trunc(b), 255} + {r, g, b, a} -> {trunc(r), trunc(g), trunc(b), trunc(a)} + end + end defp decode_points(bin, 0, _flags), do: {:ok, [], bin} diff --git a/lib/ex_codecs/spatial/codec/gsplat.ex b/lib/ex_codecs/spatial/codec/gsplat.ex index f885bc0..21d702d 100644 --- a/lib/ex_codecs/spatial/codec/gsplat.ex +++ b/lib/ex_codecs/spatial/codec/gsplat.ex @@ -118,6 +118,13 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do a `%Gaussian{}` * `{:error, %ExCodecs.Error{reason: :io_error}}` on file failures + ## Raises / exceptions + + Schema and non-`%Gaussian{}` element failures are returned. Invalid path terms + or non-keyword `opts` can raise `FunctionClauseError` or `ArgumentError` in + the path/file/keyword APIs. Malformed Gaussian field shapes can raise during + bitstring construction (same class of exceptions as `encode/2`). + ## Examples iex> alias ExCodecs.Spatial.{Gaussian, Codec.Gsplat} @@ -178,14 +185,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do defp schema_to_sh_rest(schema) when is_list(schema) do sh_rest = Keyword.get(schema, :sh_rest, 0) - - unknown = - schema - |> Enum.reject(fn - {:sh_rest, _} -> true - :sh_rest -> true - _ -> false - end) + unknown = Enum.reject(schema, &known_gspl_schema_entry?/1) cond do unknown != [] -> @@ -219,6 +219,10 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do )} end + defp known_gspl_schema_entry?(:sh_rest), do: true + defp known_gspl_schema_entry?({:sh_rest, _}), do: true + defp known_gspl_schema_entry?(_), do: false + defp gspl_header(flags, count, sh_rest) do <<@magic::binary, @version::little-unsigned-16, flags::little-unsigned-16, count::little-unsigned-64, sh_rest::little-unsigned-16>> diff --git a/lib/ex_codecs/spatial/codec/ply.ex b/lib/ex_codecs/spatial/codec/ply.ex index 0872b74..726c4dd 100644 --- a/lib/ex_codecs/spatial/codec/ply.ex +++ b/lib/ex_codecs/spatial/codec/ply.ex @@ -972,13 +972,10 @@ defmodule ExCodecs.Spatial.Codec.PLY do buf = acc <> data match_or_continue_header(io, buf) - :eof -> - header_eof_error(acc) - {:error, reason} -> ply_io_error(reason) - _other -> + _eof_or_empty -> header_eof_error(acc) end end @@ -1005,25 +1002,12 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end - defp header_eof_error("") do + defp header_eof_error(_acc) do + # Acc never contains a complete end_header here: that case is handled in + # match_or_continue_header/2 while chunks are still arriving. {:error, Error.new(:invalid_data, codec: :ply, message: "PLY header missing end_header")} end - defp header_eof_error(acc) do - case :binary.match(acc, "end_header") do - {pos, len} -> - header = binary_part(acc, 0, pos + len) - rest = binary_part(acc, pos + len, byte_size(acc) - pos - len) - - with {:ok, parsed} <- parse_header(header) do - {:ok, parsed, strip_leading_newlines(rest)} - end - - :nomatch -> - {:error, Error.new(:invalid_data, codec: :ply, message: "PLY header missing end_header")} - end - end - defp next_ply_stream_item({:ok, {:binary, io, _buf, _endian, parsed, _as, _stride, i}}) when i >= parsed.count do {:halt, {:done, io}} @@ -1079,16 +1063,11 @@ defmodule ExCodecs.Spatial.Codec.PLY do data when is_binary(data) and byte_size(data) > 0 -> take_exact_bytes(io, buf <> data, n) - :eof -> - {:error, - Error.new(:invalid_data, codec: :ply, message: "Binary PLY body too short")} - - data when is_binary(data) -> - {:error, - Error.new(:invalid_data, codec: :ply, message: "Binary PLY body too short")} - {:error, reason} -> ply_io_error(reason) + + _eof_or_short -> + {:error, Error.new(:invalid_data, codec: :ply, message: "Binary PLY body too short")} end end @@ -1111,7 +1090,10 @@ defmodule ExCodecs.Spatial.Codec.PLY do data when is_binary(data) and byte_size(data) > 0 -> take_ascii_vertex_line(io, buf <> data) - :eof -> + {:error, reason} -> + ply_io_error(reason) + + _eof_or_empty -> trimmed = String.trim(buf) if trimmed == "" do @@ -1123,16 +1105,6 @@ defmodule ExCodecs.Spatial.Codec.PLY do else {:ok, trimmed, ""} end - - {:error, reason} -> - ply_io_error(reason) - - _other -> - {:error, - Error.new(:invalid_data, - codec: :ply, - message: "Expected more ASCII vertices" - )} end end diff --git a/lib/ex_codecs/spatial/stream.ex b/lib/ex_codecs/spatial/stream.ex index 7a071ba..4d286f1 100644 --- a/lib/ex_codecs/spatial/stream.ex +++ b/lib/ex_codecs/spatial/stream.ex @@ -247,32 +247,35 @@ defmodule ExCodecs.Spatial.Stream do defp stream_encode_to_file_if_schema(data, path, opts) do if Keyword.has_key?(opts, :schema) do - case Keyword.get(opts, :format) do - :spatial_binary -> - Binary.stream_encode_to_file(enumerable_points(data), path, opts) - - :gsplat -> - Gsplat.stream_encode_to_file(enumerable_gaussians(data), path, opts) - - other when other in [:ply, nil] -> - {:error, - Error.new(:invalid_options, - message: - "encode_to_file schema streaming requires format: :spatial_binary or :gsplat" - )} - - other -> - {:error, - Error.new(:unsupported_codec, - codec: other, - message: "Unsupported spatial stream format: #{inspect(other)}" - )} - end + stream_encode_by_format(Keyword.get(opts, :format), data, path, opts) else :not_streaming end end + defp stream_encode_by_format(:spatial_binary, data, path, opts) do + Binary.stream_encode_to_file(enumerable_points(data), path, opts) + end + + defp stream_encode_by_format(:gsplat, data, path, opts) do + Gsplat.stream_encode_to_file(enumerable_gaussians(data), path, opts) + end + + defp stream_encode_by_format(other, _data, _path, _opts) when other in [:ply, nil] do + {:error, + Error.new(:invalid_options, + message: "encode_to_file schema streaming requires format: :spatial_binary or :gsplat" + )} + end + + defp stream_encode_by_format(other, _data, _path, _opts) do + {:error, + Error.new(:unsupported_codec, + codec: other, + message: "Unsupported spatial stream format: #{inspect(other)}" + )} + end + defp enumerable_points(%PointCloud{points: points}), do: points defp enumerable_points(enumerable), do: enumerable diff --git a/test/ex_codecs/spatial/stream_coverage_test.exs b/test/ex_codecs/spatial/stream_coverage_test.exs new file mode 100644 index 0000000..af2fc9b --- /dev/null +++ b/test/ex_codecs/spatial/stream_coverage_test.exs @@ -0,0 +1,461 @@ +defmodule ExCodecs.Spatial.StreamCoverageTest do + use ExUnit.Case, async: true + + alias ExCodecs.Spatial + alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} + alias ExCodecs.Spatial.Codec.{Binary, Gsplat, PLY} + alias ExCodecs.Spatial.Stream, as: SpatialStream + + defp tmp(ext) do + Path.join( + System.tmp_dir!(), + "ex_codecs_stream_cov_#{System.unique_integer([:positive])}#{ext}" + ) + end + + describe "Binary stream_encode_to_file edges" do + test "rejects non-Point elements" do + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + + assert {:error, %{reason: :invalid_data, codec: :spatial_binary}} = + Binary.stream_encode_to_file([Point.new(0, 0, 0), :nope], path, schema: []) + end + + test "accepts map schema, keyword schema, alpha, and nil color fill" do + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + + assert :ok = + Binary.stream_encode_to_file([Point.new(1, 2, 3)], path, + schema: %{color: true, normal: false} + ) + + assert :ok = + Binary.stream_encode_to_file( + [Point.new(1, 2, 3, color: {1, 2, 3, 4}, normal: {0, 1, 0})], + path, + schema: [color: true, alpha: true, normal: true] + ) + + assert {:ok, decoded} = Binary.decode(File.read!(path)) + assert hd(decoded.points).color == {1, 2, 3, 4} + + assert :ok = + Binary.stream_encode_to_file([Point.new(9, 8, 7)], path, schema: [:color, :alpha]) + + assert {:ok, filled} = Binary.decode(File.read!(path)) + assert hd(filled.points).color == {0, 0, 0, 255} + end + + test "rejects unknown and non-list/map schemas and unwritable paths" do + assert {:error, %{reason: :invalid_options}} = + Binary.stream_encode_to_file([Point.new(0, 0, 0)], tmp(".excp"), schema: [:bogus]) + + assert {:error, %{reason: :invalid_options}} = + Binary.stream_encode_to_file([Point.new(0, 0, 0)], tmp(".excp"), schema: [foo: true]) + + assert {:error, %{reason: :invalid_options}} = + Binary.stream_encode_to_file([Point.new(0, 0, 0)], tmp(".excp"), schema: :color) + + assert {:error, %{reason: :io_error}} = + Binary.stream_encode_to_file([Point.new(0, 0, 0)], "/no/such/dir/x.excp", schema: []) + end + end + + describe "Binary stream_decode file edges" do + test "unsupported version, empty file, short header, eof after full record" do + bad_ver = tmp(".excp") + empty = tmp(".excp") + short = tmp(".excp") + eof_mid = tmp(".excp") + on_exit(fn -> Enum.each([bad_ver, empty, short, eof_mid], &File.rm/1) end) + + File.write!( + bad_ver, + <<"EXCP", 99::little-unsigned-16, 0::little-unsigned-16, 1::little-unsigned-64>> + ) + + File.write!(empty, "") + File.write!(short, "EXCP") + + # count=2 but only one xyz record → second read hits :eof + one = <<1.0::little-float-32, 2.0::little-float-32, 3.0::little-float-32>> + + File.write!( + eof_mid, + <<"EXCP", 1::little-unsigned-16, 0::little-unsigned-16, 2::little-unsigned-64, one::binary>> + ) + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(bad_ver, source: :file) |> Enum.to_list() + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(empty, source: :file) |> Enum.to_list() + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(short, source: :file) |> Enum.to_list() + + items = Binary.stream_decode(eof_mid, source: :file) |> Enum.to_list() + assert match?([%Point{}, {:error, %{reason: :invalid_data}}], items) + end + + test "early halt closes open file handle" do + cloud = PointCloud.new([Point.new(1, 2, 3), Point.new(4, 5, 6)]) + {:ok, bin} = Binary.encode(cloud) + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + assert [%Point{}] = + Binary.stream_decode(path, source: :file) |> Stream.take(1) |> Enum.to_list() + end + + test "source: :binary and alpha stride decode from file" do + cloud = PointCloud.new([Point.new(1, 2, 3, color: {1, 2, 3, 4})]) + {:ok, bin} = Binary.encode(cloud) + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + assert [%Point{color: {1, 2, 3, 4}}] = + Binary.stream_decode(bin, source: :binary) |> Enum.to_list() + + assert [%Point{color: {1, 2, 3, 4}}] = + Binary.stream_decode(path, source: :file) |> Enum.to_list() + end + + test "schema color without alpha zero-fills missing RGB and keeps RGB triples" do + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + + assert :ok = + Binary.stream_encode_to_file([Point.new(1, 2, 3)], path, schema: [:color]) + + assert {:ok, decoded} = Binary.decode(File.read!(path)) + assert hd(decoded.points).color == {0, 0, 0} + + assert :ok = + Binary.stream_encode_to_file( + [Point.new(1, 2, 3, color: {9, 8, 7}), Point.new(0, 0, 0, color: {1, 2, 3, 4})], + path, + schema: [:color] + ) + + assert {:ok, colored} = Binary.decode(File.read!(path)) + assert Enum.map(colored.points, & &1.color) == [{9, 8, 7}, {1, 2, 3}] + end + + test "schema keyword tuples cover known and unknown keys" do + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + + assert :ok = + Binary.stream_encode_to_file([Point.new(1, 2, 3)], path, + schema: [{:color, true}, {:normal, false}] + ) + + assert {:error, %{reason: :invalid_options}} = + Binary.stream_encode_to_file([Point.new(1, 2, 3)], path, + schema: [{:color, true}, {:wat, true}] + ) + end + end + + describe "Gsplat stream_encode_to_file edges" do + test "rejects non-Gaussian elements and bad schemas" do + path = tmp(".gspl") + on_exit(fn -> File.rm(path) end) + + assert {:error, %{reason: :invalid_data, codec: :gsplat}} = + Gsplat.stream_encode_to_file([Gaussian.new({0, 0, 0}), :nope], path, schema: []) + + assert {:error, %{reason: :invalid_options}} = + Gsplat.stream_encode_to_file([Gaussian.new({0, 0, 0})], path, schema: [:bogus]) + + assert {:error, %{reason: :invalid_options}} = + Gsplat.stream_encode_to_file([Gaussian.new({0, 0, 0})], path, schema: [sh_rest: -1]) + + assert {:error, %{reason: :invalid_options}} = + Gsplat.stream_encode_to_file([Gaussian.new({0, 0, 0})], path, schema: "nope") + + assert {:error, %{reason: :io_error}} = + Gsplat.stream_encode_to_file([Gaussian.new({0, 0, 0})], "/no/such/dir/x.gspl", + schema: [] + ) + end + + test "map schema and sh_rest round-trip" do + path = tmp(".gspl") + on_exit(fn -> File.rm(path) end) + + g = + Gaussian.new({1.0, 2.0, 3.0}, + color: {0.1, 0.2, 0.3}, + sh: [[0.1, 0.2, 0.3], [0.4, 0.5, 0.6]] + ) + + assert :ok = Gsplat.stream_encode_to_file([g], path, schema: %{sh_rest: 3}) + assert {:ok, decoded} = Gsplat.decode(File.read!(path)) + assert length(decoded.gaussians) == 1 + + assert :ok = Gsplat.stream_encode_to_file([g], path, schema: [:sh_rest]) + end + end + + describe "Gsplat stream_decode file edges" do + test "unsupported version, empty, short header, eof after full record" do + bad_ver = tmp(".gspl") + empty = tmp(".gspl") + short = tmp(".gspl") + eof_mid = tmp(".gspl") + on_exit(fn -> Enum.each([bad_ver, empty, short, eof_mid], &File.rm/1) end) + + File.write!( + bad_ver, + <<"GSPL", 9::little-unsigned-16, 0::little-unsigned-16, 1::little-unsigned-64, + 0::little-unsigned-16>> + ) + + File.write!(empty, "") + File.write!(short, "GSPL") + + zeros = :binary.copy(<<0>>, 56) + + File.write!( + eof_mid, + <<"GSPL", 1::little-unsigned-16, 0::little-unsigned-16, 2::little-unsigned-64, + 0::little-unsigned-16, zeros::binary>> + ) + + assert [{:error, %{reason: :invalid_data}}] = + Gsplat.stream_decode(bad_ver, source: :file) |> Enum.to_list() + + assert [{:error, %{reason: :invalid_data}}] = + Gsplat.stream_decode(empty, source: :file) |> Enum.to_list() + + assert [{:error, %{reason: :invalid_data}}] = + Gsplat.stream_decode(short, source: :file) |> Enum.to_list() + + items = Gsplat.stream_decode(eof_mid, source: :file) |> Enum.to_list() + assert match?([%Gaussian{}, {:error, %{reason: :invalid_data}}], items) + end + + test "early halt and source: :binary" do + cloud = GaussianCloud.new([Gaussian.new({0, 0, 0}), Gaussian.new({1, 1, 1})]) + {:ok, bin} = Gsplat.encode(cloud) + path = tmp(".gspl") + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + assert [%Gaussian{}] = + Gsplat.stream_decode(path, source: :file) |> Stream.take(1) |> Enum.to_list() + + assert [%Gaussian{}, %Gaussian{}] = + Gsplat.stream_decode(bin, source: :binary) |> Enum.to_list() + end + end + + describe "PLY stream_decode file edges" do + test "truncated ascii body and blank lines" do + path = tmp(".ply") + on_exit(fn -> File.rm(path) end) + + File.write!(path, """ + ply + format ascii 1.0 + element vertex 2 + property float x + property float y + property float z + end_header + 1.0 2.0 3.0 + + """) + + items = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert match?([%Point{}, {:error, %{reason: :invalid_data}}], items) + end + + test "ascii early halt closes open handle" do + cloud = PointCloud.new([Point.new(1, 2, 3), Point.new(4, 5, 6)]) + {:ok, bin} = PLY.encode(cloud, format: :ascii) + path = tmp(".ply") + on_exit(fn -> File.rm(path) end) + File.write!(path, bin) + + assert [%Point{}] = + PLY.stream_decode(path, source: :file) |> Stream.take(1) |> Enum.to_list() + end + + test "ascii last vertex without trailing newline" do + path = tmp(".ply") + on_exit(fn -> File.rm(path) end) + + File.write!( + path, + "ply\nformat ascii 1.0\nelement vertex 1\nproperty float x\nproperty float y\nproperty float z\nend_header\n1.5 2.5 3.5" + ) + + assert [%Point{x: x}] = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert_in_delta x, 1.5, 0.0001 + end + + test "corrupt header after end_header marker and oversized header" do + bad = tmp(".ply") + huge = tmp(".ply") + on_exit(fn -> Enum.each([bad, huge], &File.rm/1) end) + + File.write!(bad, "notply\nend_header\n") + + assert [{:error, %{reason: :invalid_data}}] = + PLY.stream_decode(bad, source: :file) |> Enum.to_list() + + # > 1 MiB without end_header + File.write!(huge, ["ply\nformat ascii 1.0\n"] ++ List.duplicate("comment pad\n", 90_000)) + + assert [{:error, %{reason: :invalid_data, message: message}}] = + PLY.stream_decode(huge, source: :file) |> Enum.to_list() + + assert message =~ "exceeds" + end + + test "empty file, early halt, and binary_be stream" do + empty = tmp(".ply") + be = tmp(".ply") + on_exit(fn -> Enum.each([empty, be], &File.rm/1) end) + + File.write!(empty, "") + + assert [{:error, %{reason: :invalid_data}}] = + PLY.stream_decode(empty, source: :file) |> Enum.to_list() + + cloud = PointCloud.new([Point.new(1, 2, 3), Point.new(4, 5, 6)]) + {:ok, bin} = PLY.encode(cloud, ply_format: :binary_be) + File.write!(be, bin) + + assert [%Point{}] = + PLY.stream_decode(be, source: :file) |> Stream.take(1) |> Enum.to_list() + + assert [%Point{}, %Point{}] = + PLY.stream_decode(be, source: :file) |> Enum.to_list() + end + + test "missing end_header at eof" do + path = tmp(".ply") + on_exit(fn -> File.rm(path) end) + File.write!(path, "ply\nformat ascii 1.0\nelement vertex 0\n") + + assert [{:error, %{reason: :invalid_data}}] = + PLY.stream_decode(path, source: :file) |> Enum.to_list() + end + + test "binary body split across header-scan chunk boundary" do + path = tmp(".ply") + on_exit(fn -> File.rm(path) end) + + # First 4096-byte read ends with end_header + 4 body bytes (< 12-byte stride); + # the rest of the vertex is in the next read (covers take_exact_bytes refill). + header = """ + ply + format binary_little_endian 1.0 + element vertex 1 + property float x + property float y + property float z + end_header + """ + + body = <<1.0::little-float-32, 2.0::little-float-32, 3.0::little-float-32>> + pad_len = 4096 - byte_size(header) - 4 + assert pad_len > 0 + # comments must appear before end_header — pad inside header instead + comment = "comment " <> String.duplicate("p", 60) <> "\n" + + base = + "ply\nformat binary_little_endian 1.0\nelement vertex 1\nproperty float x\nproperty float y\nproperty float z\n" + + suffix = "end_header\n" + need = 4096 - 4 - byte_size(base) - byte_size(suffix) + comments = String.duplicate(comment, div(need, byte_size(comment)) + 2) + header2 = base <> binary_part(comments, 0, need) <> suffix + assert byte_size(header2) == 4096 - 4 + + File.write!(path, header2 <> body) + + assert [%Point{x: x}] = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert_in_delta x, 1.0, 0.0001 + end + + test "ascii vertex line split across read chunks" do + path = tmp(".ply") + on_exit(fn -> File.rm(path) end) + + base = """ + ply + format ascii 1.0 + element vertex 1 + property float x + property float y + property float z + """ + + suffix = "end_header\n" + # Leave "1.0 2.0 3" in the first 4096 chunk and "0\n" in the next. + partial = "1.0 2.0 3" + rest = "0\n" + need = 4096 - byte_size(base) - byte_size(suffix) - byte_size(partial) + comments = String.duplicate("comment padline\n", div(need, 16) + 2) + header = base <> binary_part(comments, 0, need) <> suffix + assert byte_size(header <> partial) == 4096 + + File.write!(path, header <> partial <> rest) + + assert [%Point{z: z}] = PLY.stream_decode(path, source: :file) |> Enum.to_list() + assert_in_delta z, 30.0, 0.0001 + end + end + + describe "Spatial.Stream encode_to_file schema routing" do + test "streams GSPL from GaussianCloud and rejects unsupported format" do + path = tmp(".gspl") + list_path = tmp(".gspl") + on_exit(fn -> Enum.each([path, list_path], &File.rm/1) end) + cloud = %GaussianCloud{gaussians: [Gaussian.new({1.0, 2.0, 3.0})]} + + assert :ok = + SpatialStream.encode_to_file(cloud, path, format: :gsplat, schema: [sh_rest: 0]) + + assert [%Gaussian{}] = + Spatial.stream_decode(path, format: :gsplat, source: :file) |> Enum.to_list() + + assert :ok = + SpatialStream.encode_to_file([Gaussian.new({0.0, 1.0, 2.0})], list_path, + format: :gsplat, + schema: [] + ) + + assert {:error, %{reason: :unsupported_codec}} = + SpatialStream.encode_to_file([Point.new(0, 0, 0)], tmp(".bin"), + format: :sog, + schema: [] + ) + end + + test "streams EXCP from PointCloud" do + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + cloud = PointCloud.new([Point.new(1, 2, 3, color: {9, 8, 7})]) + + assert :ok = + SpatialStream.encode_to_file(cloud, path, + format: :spatial_binary, + schema: [:color] + ) + + assert [%Point{color: {9, 8, 7}}] = + Spatial.stream_decode(path, format: :spatial_binary, source: :file) + |> Enum.to_list() + end + end +end From a265a6861873f3c50eedc6eb85cb89e1ffc52d9a Mon Sep 17 00:00:00 2001 From: thanos Date: Sat, 18 Jul 2026 08:01:23 -0400 Subject: [PATCH 3/6] 0.2.3 release prep is in place. Summary of updates: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Version mix.exs → 0.2.3 native/ex_codecs_native/Cargo.toml → 0.2.3 Changelog Moved Unreleased items into [0.2.3] - 2026-07-18 (streaming I/O, schema encode-to-file, Rust accel, Mix cli/0) Docs README install pin, streaming limitations, native layer note docs/spatial_formats.md, docs/architecture.md, guides/understanding_spatial_codecs.md All livebooks → {:ex_codecs, "~> 0.2.3"}; spatial livebook streaming section updated --- CHANGELOG.md | 29 +- README.md | 12 +- docs/architecture.md | 15 +- docs/spatial_formats.md | 16 +- guides/understanding_spatial_codecs.md | 9 +- lib/ex_codecs/native.ex | 39 ++ lib/ex_codecs/spatial/accel.ex | 243 +++++++ lib/ex_codecs/spatial/codec/binary.ex | 342 ++++++++-- lib/ex_codecs/spatial/codec/gsplat.ex | 333 +++++++++- lib/ex_codecs/spatial/codec/ply.ex | 175 ++++- livebooks/01_introduction.livemd | 2 +- livebooks/02_compression_fundamentals.livemd | 2 +- livebooks/03_codec_comparison.livemd | 2 +- livebooks/04_building_storage_systems.livemd | 2 +- livebooks/05_zarr_style_workloads.livemd | 2 +- livebooks/06_spatial_codecs.livemd | 9 +- mix.exs | 11 +- native/ex_codecs_native/Cargo.lock | 12 +- native/ex_codecs_native/Cargo.toml | 3 +- native/ex_codecs_native/src/lib.rs | 2 + native/ex_codecs_native/src/spatial.rs | 621 ++++++++++++++++++ .../ex_codecs/spatial/accel_coverage_test.exs | 551 ++++++++++++++++ .../ex_codecs/spatial/accel_property_test.exs | 213 ++++++ .../spatial/stream_coverage_test.exs | 2 +- 24 files changed, 2513 insertions(+), 134 deletions(-) create mode 100644 lib/ex_codecs/spatial/accel.ex create mode 100644 native/ex_codecs_native/src/spatial.rs create mode 100644 test/ex_codecs/spatial/accel_coverage_test.exs create mode 100644 test/ex_codecs/spatial/accel_property_test.exs diff --git a/CHANGELOG.md b/CHANGELOG.md index a76a3a2..1bcc839 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,20 +7,35 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] -### Changed - -- EXCP (`:spatial_binary`), GSPL (`:gsplat`), and PLY `stream_decode` with - `source: :file` (or `:auto` path detection) now read the header and then - **one record/vertex at a time** from disk (bounded memory). Binary PLY uses - a fixed stride; ASCII PLY reads lines. In-memory binaries still materialize; - a future Rust backend may memory-map those. +## [0.2.3] - 2026-07-18 ### Added +- **Incremental spatial file I/O** — EXCP (`:spatial_binary`), GSPL (`:gsplat`), + and PLY `stream_decode` with `source: :file` (or `:auto` path detection) read + the header then **one record/vertex at a time** from disk (bounded memory). + Binary PLY uses a fixed stride; ASCII PLY reads lines. - `Binary.stream_encode_to_file/3` and `Gsplat.stream_encode_to_file/3` — incremental file writes with an explicit `:schema` (placeholder header, seek-back count). `Spatial.Stream.encode_to_file/3` uses these when `:schema` is present with `format: :spatial_binary` or `:gsplat`. +- **Spatial Rust acceleration** (DirtyCpu NIFs): chunked EXCP/GSPL pack & + unpack, mmap-backed file `stream_decode`, binary PLY body unpack, and + chunked `stream_encode_to_file`. Pass `accel: false` to force pure Elixir. + Property tests compare both backends byte-for-byte / structurally. + +### Changed + +- Mix `preferred_cli_env` moved into `cli/0` (`preferred_envs`) for Mix 1.20+. +- In-memory spatial `stream_decode` uses chunked Rust unpack when the spatial + NIF is loaded; otherwise it still materializes through `decode/2`. + +### Notes + +- Wire layouts for EXCP / GSPL / PLY remain the **v0.2.0 freeze** in + `docs/spatial_formats.md` (Rust output is byte-compatible). +- Precompiled NIF checksums must be regenerated when publishing GitHub release + artifacts for `0.2.3`. ## [0.2.0] - 2026-07-16 diff --git a/README.md b/README.md index 5336218..92dac41 100644 --- a/README.md +++ b/README.md @@ -42,7 +42,7 @@ Add `ex_codecs` to your list of dependencies in `mix.exs`: ```elixir def deps do [ - {:ex_codecs, "~> 0.2.0"} + {:ex_codecs, "~> 0.2.3"} ] end ``` @@ -180,7 +180,10 @@ and the frozen [Spatial wire formats](https://hexdocs.pm/ex_codecs/spatial_forma bytes/ratios are not guaranteed identical to C libzstd. - **Decompression** defaults to a **256 MiB** `max_output_size`. Raise it only for trusted inputs; do not decompress untrusted payloads without a tight limit. -- Spatial `stream_*` helpers **materialize** full payloads today. +- Spatial **file** `stream_decode` is incremental (bounded memory). In-memory + binaries use chunked Rust unpack when available, otherwise materialize. + `stream_encode/2` still collects the enumerable, then encodes once; use + `encode_to_file/3` with an explicit `:schema` for EXCP/GSPL file streaming. ## Architecture @@ -196,8 +199,9 @@ ExCodecs is layered as follows: 3. **Shared codec catalog** (`ExCodecs.CodecRegistry`) — ETS map of codec atoms to modules, categories, interface shapes, and metadata, populated at startup. -4. **Native NIFs** (`ExCodecs.Native`) — pure-Rust compression via - `rustler_precompiled` (or local compile). +4. **Native NIFs** (`ExCodecs.Native`) — pure-Rust compression plus optional + spatial DirtyCpu / mmap acceleration via `rustler_precompiled` (or local + compile). 5. **Category discovery** — `available_codecs/0` lists the whole catalog; `available_codecs/1` filters it, and diff --git a/docs/architecture.md b/docs/architecture.md index 7118de7..51825c5 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -97,7 +97,8 @@ The implementation is layered in four tiers: modules expose the category's struct↔format contract. 4. **Native layer** -- Compression modules delegate to the Rustler NIF - (`ExCodecs.Native`). Spatial codecs are Elixir implementations in v0.2.0. + (`ExCodecs.Native`). Spatial codecs keep a pure-Elixir path and, since + v0.2.3, optional DirtyCpu / mmap acceleration via the same NIF crate. --- @@ -727,12 +728,14 @@ predictable and navigable. ## Codec Categories -### Spatial (implemented in 0.2.0) +### Spatial (since 0.2.0; Rust accel in 0.2.3) -Spatial codecs map structured geometric types to interchange formats. They are -pure Elixir and use the specialized `ExCodecs.Spatial` API rather than forcing -structs through the binary-only `ExCodecs.Codec` callbacks. They are registered -in the shared catalog with `category: :spatial` and `interface: :spatial`. +Spatial codecs map structured geometric types to interchange formats. They use +the specialized `ExCodecs.Spatial` API rather than forcing structs through the +binary-only `ExCodecs.Codec` callbacks. They are registered in the shared +catalog with `category: :spatial` and `interface: :spatial`. Hot paths may use +DirtyCpu pack/unpack and mmap-backed file streams when the NIF is loaded +(`accel: false` forces Elixir). `ExCodecs.available_codecs/0` lists all available entries, `ExCodecs.available_codecs(:spatial)` filters the shared catalog, and diff --git a/docs/spatial_formats.md b/docs/spatial_formats.md index d1beb1f..e091e36 100644 --- a/docs/spatial_formats.md +++ b/docs/spatial_formats.md @@ -1,8 +1,9 @@ # Spatial wire formats (v0.2.0 freeze) -This document freezes the on-wire layouts used by `ExCodecs.Spatial` in **v0.2.0**. -A future Rust backend (v0.2.1) must produce **byte-compatible** output for these -formats. There is **no CRC / checksum** in v1; integrity is the caller's responsibility. +This document freezes the on-wire layouts used by `ExCodecs.Spatial` since +**v0.2.0**. The Rust spatial acceleration shipped in **v0.2.3** must produce +**byte-compatible** output for these formats (verified by property tests). +There is **no CRC / checksum** in v1; integrity is the caller's responsibility. Attribute keys on points are **strings** after construction / decode. Prefer string keys when building clouds for cross-backend compatibility. @@ -45,13 +46,16 @@ normal `{0,0,0}`). Mixed optional fields therefore round-trip with defaults fill - **EXCP / GSPL / PLY + `source: :file`** (or `:auto` path detection): header then one record/vertex at a time from disk (bounded memory). Binary PLY uses a fixed property stride; ASCII PLY reads lines. -- In-memory binaries still **materialize** through `decode/2`, then enumerate. +- In-memory binaries use **chunked Rust unpack** when the spatial NIF is + loaded; otherwise they materialize through `decode/2`. - `stream_encode` still collects the enumerable, then encodes once. - `encode_to_file/3` with an explicit `:schema` streams EXCP/GSPL to disk - (placeholder header + seek-back count). Without `:schema`, encode then write. + (placeholder header + seek-back count; chunked Rust pack when available). + Without `:schema`, encode then write. +- Pass `accel: false` on encode/decode/stream helpers to force pure Elixir. Prefer explicit `source: :file` or `source: :binary` when the argument is -ambiguous. A future Rust backend may memory-map large in-memory binaries. +ambiguous. File streams prefer mmap + DirtyCpu unpack when Accel is available. `:auto` treats a binary as a path only when it looks path-like (under 4 KiB, no `ply`/`EXCP`/`GSPL` magic, and contains `/` or `\` or ends with `.ply`/`.excp`/ diff --git a/guides/understanding_spatial_codecs.md b/guides/understanding_spatial_codecs.md index ea09324..8699e70 100644 --- a/guides/understanding_spatial_codecs.md +++ b/guides/understanding_spatial_codecs.md @@ -368,9 +368,12 @@ clouds. Magic bytes: `"GSPL"`. Wire layouts and schema rules are frozen in [Spatial wire formats](../docs/spatial_formats.md). -Stream helpers: EXCP/GSPL **files** decode record-by-record from disk; PLY and -in-memory binaries still materialize. Prefer explicit `source: :file` or -`source: :binary`. +Stream helpers (v0.2.3+): EXCP / GSPL / PLY **files** decode record-by-record +from disk (bounded memory; mmap + Rust unpack when available). In-memory +binaries use chunked Rust unpack when the NIF is loaded, otherwise materialize. +`stream_encode/2` still collects then encodes; use `encode_to_file/3` with an +explicit `:schema` for EXCP/GSPL incremental writes. Prefer explicit +`source: :file` or `source: :binary`. ## Why a separate `ExCodecs.Spatial` API? diff --git a/lib/ex_codecs/native.ex b/lib/ex_codecs/native.ex index 39bbb06..bb24842 100644 --- a/lib/ex_codecs/native.ex +++ b/lib/ex_codecs/native.ex @@ -65,6 +65,45 @@ defmodule ExCodecs.Native do @doc false def codec_versions, do: :erlang.nif_error(:nif_not_loaded) + # --- Spatial acceleration (DirtyCpu) --------------------------------------- + + @doc false + def excp_unpack(_data, _flags, _offset, _max_count), do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def excp_pack(_records, _flags), do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def gspl_unpack(_data, _sh_rest, _offset, _max_count), do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def gspl_pack(_records, _sh_rest), do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def ply_binary_unpack(_data, _types, _little_endian, _offset, _max_count), + do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def spatial_mmap_open(_path), do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def spatial_mmap_len(_resource), do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def excp_unpack_mmap(_resource, _flags, _offset, _max_count), + do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def gspl_unpack_mmap(_resource, _sh_rest, _offset, _max_count), + do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def ply_binary_unpack_mmap(_resource, _types, _little_endian, _offset, _max_count), + do: :erlang.nif_error(:nif_not_loaded) + + @doc false + def spatial_append_file(_path, _data), do: :erlang.nif_error(:nif_not_loaded) + @doc """ Returns `true` when the native NIF library is loaded. diff --git a/lib/ex_codecs/spatial/accel.ex b/lib/ex_codecs/spatial/accel.ex new file mode 100644 index 0000000..3e38fb0 --- /dev/null +++ b/lib/ex_codecs/spatial/accel.ex @@ -0,0 +1,243 @@ +defmodule ExCodecs.Spatial.Accel do + @moduledoc false + + # Thin facade over DirtyCpu spatial NIFs. Callers fall back to pure Elixir + # when `available?/0` is false or a call returns `:nif_not_loaded`. + + alias ExCodecs.Native + alias ExCodecs.Spatial.{Gaussian, Point} + + @default_chunk 4_096 + + @ply_type_tags %{ + char: 1, + uchar: 2, + short: 3, + ushort: 4, + int: 5, + uint: 6, + float: 7, + double: 8 + } + + def available? do + case safe(fn -> Native.codec_versions() end) do + map when is_map(map) -> + Map.has_key?(map, "spatial") or Map.has_key?(map, :spatial) + + # coveralls-ignore-start + _ -> + false + # coveralls-ignore-stop + end + end + + def chunk_size, do: @default_chunk + + # --- EXCP ----------------------------------------------------------------- + + def excp_unpack(data, flags, offset \\ 0, max_count \\ @default_chunk) + when is_binary(data) and is_integer(flags) do + nif_chunk(fn -> Native.excp_unpack(data, flags, offset, max_count) end, &rows_to_points/1) + end + + def excp_unpack_mmap(resource, flags, offset \\ 0, max_count \\ @default_chunk) do + nif_chunk( + fn -> Native.excp_unpack_mmap(resource, flags, offset, max_count) end, + &rows_to_points/1 + ) + end + + def excp_pack(points, flags) when is_list(points) and is_integer(flags) do + records = Enum.map(points, &point_to_row/1) + nif_binary(fn -> Native.excp_pack(records, flags) end) + rescue + _ in [FunctionClauseError, MatchError] -> {:error, :invalid_data} + end + + # --- GSPL ----------------------------------------------------------------- + + def gspl_unpack(data, sh_rest, offset \\ 0, max_count \\ @default_chunk) + when is_binary(data) and is_integer(sh_rest) do + nif_chunk( + fn -> Native.gspl_unpack(data, sh_rest, offset, max_count) end, + &rows_to_gaussians(&1, sh_rest) + ) + end + + def gspl_unpack_mmap(resource, sh_rest, offset \\ 0, max_count \\ @default_chunk) do + nif_chunk( + fn -> Native.gspl_unpack_mmap(resource, sh_rest, offset, max_count) end, + &rows_to_gaussians(&1, sh_rest) + ) + end + + def gspl_pack(gaussians, sh_rest) when is_list(gaussians) and is_integer(sh_rest) do + records = Enum.map(gaussians, &gaussian_to_row(&1, sh_rest)) + nif_binary(fn -> Native.gspl_pack(records, sh_rest) end) + rescue + _ in [FunctionClauseError, MatchError] -> {:error, :invalid_data} + end + + # --- Binary PLY ----------------------------------------------------------- + + def ply_type_tag(type) when is_atom(type), do: Map.fetch!(@ply_type_tags, type) + + def ply_binary_unpack(data, types, endian, offset \\ 0, max_count \\ @default_chunk) + when is_binary(data) and is_list(types) do + tags = Enum.map(types, &ply_type_tag/1) + little? = endian in [:binary_le, :little, true] + + nif_chunk( + fn -> Native.ply_binary_unpack(data, tags, little?, offset, max_count) end, + & &1 + ) + end + + def ply_binary_unpack_mmap(resource, types, endian, offset \\ 0, max_count \\ @default_chunk) do + tags = Enum.map(types, &ply_type_tag/1) + little? = endian in [:binary_le, :little, true] + + nif_chunk( + fn -> Native.ply_binary_unpack_mmap(resource, tags, little?, offset, max_count) end, + & &1 + ) + end + + # --- mmap ----------------------------------------------------------------- + + def mmap_open(path) when is_binary(path) do + case safe(fn -> Native.spatial_mmap_open(path) end) do + {:ok, resource} -> {:ok, resource} + {:error, _} = err -> err + end + end + + def mmap_len(resource) do + case safe(fn -> Native.spatial_mmap_len(resource) end) do + n when is_integer(n) -> + {:ok, n} + + {:error, _} = err -> + err + + # coveralls-ignore-start + other -> + {:error, {:unexpected, other}} + # coveralls-ignore-stop + end + end + + def append_file(path, data) when is_binary(path) and is_binary(data) do + case safe(fn -> Native.spatial_append_file(path, data) end) do + :ok -> :ok + {:error, _} = err -> err + end + end + + # --- row codecs ----------------------------------------------------------- + + def point_to_row(%Point{} = p) do + {p.x, p.y, p.z, p.color, p.normal} + end + + def row_to_point({x, y, z, color, normal}) do + Point.new(x, y, z, color: color, normal: normal) + end + + def gaussian_to_row(%Gaussian{} = g, sh_rest) do + {x, y, z} = g.position + {r, gc, b} = g.color + {sx, sy, sz} = g.scale + {rw, rx, ry, rz} = g.rotation + + sh = + case g.sh do + nil -> [] + [_dc | rest] -> List.flatten(rest) + list when is_list(list) -> list |> List.flatten() |> Enum.drop(3) + end + |> then(fn vals -> + vals + |> Stream.concat(Stream.cycle([0.0])) + |> Enum.take(sh_rest) + end) + + {{x, y, z}, {r, gc, b}, g.opacity, {sx, sy, sz}, {rw, rx, ry, rz}, sh} + end + + def row_to_gaussian( + {{x, y, z}, {r, gc, b}, opacity, {sx, sy, sz}, {rw, rx, ry, rz}, sh_vals}, + sh_rest + ) do + sh = + if sh_rest == 0 do + nil + else + [[r, gc, b] | Enum.chunk_every(sh_vals, 3)] + end + + Gaussian.new({x, y, z}, + color: {r, gc, b}, + opacity: opacity, + scale: {sx, sy, sz}, + rotation: {rw, rx, ry, rz}, + sh: sh + ) + end + + defp rows_to_points(rows), do: Enum.map(rows, &row_to_point/1) + + defp rows_to_gaussians(rows, sh_rest), + do: Enum.map(rows, &row_to_gaussian(&1, sh_rest)) + + defp nif_chunk(fun, map_rows) do + case safe(fun) do + {:ok, {rows, next_offset}} when is_list(rows) and is_integer(next_offset) -> + {:ok, {map_rows.(rows), next_offset}} + + {:error, _} = err -> + err + + # coveralls-ignore-start + other -> + {:error, {:unexpected, other}} + # coveralls-ignore-stop + end + end + + defp nif_binary(fun) do + case safe(fun) do + {:ok, bin} when is_binary(bin) -> + {:ok, bin} + + {:error, _} = err -> + err + + # coveralls-ignore-start + other -> + {:error, {:unexpected, other}} + # coveralls-ignore-stop + end + end + + defp safe(fun) do + fun.() + rescue + e -> + # Rustler may raise ArgumentError for bad resources; that is not always an + # `%ErlangError{}`, so discriminate on the struct rather than `rescue in`. + case e do + # coveralls-ignore-start + %ErlangError{original: :nif_not_loaded} -> + {:error, :nif_not_loaded} + + %ErlangError{original: other} -> + {:error, other} + + # coveralls-ignore-stop + other -> + {:error, Exception.message(other)} + end + end +end diff --git a/lib/ex_codecs/spatial/codec/binary.ex b/lib/ex_codecs/spatial/codec/binary.ex index d8744cd..b5ee79d 100644 --- a/lib/ex_codecs/spatial/codec/binary.ex +++ b/lib/ex_codecs/spatial/codec/binary.ex @@ -17,6 +17,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do """ alias ExCodecs.Error + alias ExCodecs.Spatial.Accel alias ExCodecs.Spatial.{Metadata, Point, PointCloud} @magic "EXCP" @@ -72,7 +73,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do @spec encode(PointCloud.t(), keyword()) :: {:ok, binary()} | {:error, Error.t()} def encode(data, opts \\ []) - def encode(%PointCloud{points: points}, _opts) do + def encode(%PointCloud{points: points}, opts) do has_color? = Enum.any?(points, &Point.colored?/1) has_alpha? = Enum.any?(points, fn p -> match?({_, _, _, _}, p.color) end) has_normal? = Enum.any?(points, &Point.has_normal?/1) @@ -87,13 +88,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do <<@magic::binary, @version::little-unsigned-16, flags::little-unsigned-16, length(points)::little-unsigned-64>> - body = - IO.iodata_to_binary( - Enum.map(points, fn p -> - encode_point(p, flags) - end) - ) - + body = encode_points_body(points, flags, opts) {:ok, header <> body} end @@ -151,17 +146,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do {:ok, io} <- open_write(path) do try do :ok = IO.binwrite(io, excp_header(flags, 0)) - - count = - Enum.reduce(enumerable, 0, fn - %Point{} = point, n -> - :ok = IO.binwrite(io, encode_point(point, flags)) - n + 1 - - other, _n -> - throw({:bad_point, other}) - end) - + count = write_point_chunks(enumerable, io, flags, opts) {:ok, 0} = :file.position(io, 0) :ok = IO.binwrite(io, excp_header(flags, count)) :ok @@ -293,7 +278,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do def decode( <<@magic::binary, version::little-unsigned-16, flags::little-unsigned-16, count::little-unsigned-64, rest::binary>>, - _opts + opts ) do if version != @version do {:error, @@ -302,7 +287,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do message: "Unsupported binary point format version #{version}" )} else - with {:ok, points, _} <- decode_points(rest, count, flags) do + with {:ok, points, _} <- decode_points(rest, count, flags, opts) do meta = Metadata.new(entries: %{"format" => "excp", "version" => version}) {:ok, PointCloud.new(points, metadata: meta)} end @@ -322,15 +307,14 @@ defmodule ExCodecs.Spatial.Codec.Binary do ## Streaming behavior - * `source: :file` (or `:auto` when the argument looks like a path to a - regular file) reads the 16-byte header, then **one record at a time** - via `IO.binread/2`. Peak memory is O(header + one record), not the - whole cloud. - * `source: :binary` (or `:auto` for an in-memory EXCP payload) still - materializes through `decode/2`, then yields the list. A future Rust - backend may memory-map large binaries instead. + * `source: :file` (or `:auto` path detection): header then chunked + records — Rust mmap + DirtyCpu unpack when available, otherwise + `IO.binread/2` per record. + * `source: :binary`: chunked Rust unpack when available; otherwise + materializes through `decode/2`. Prefer `source: :file` for multi‑MB / multi‑GB `.excp` paths. + Pass `accel: false` to force the pure-Elixir path. ## Arguments @@ -370,7 +354,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do {:ok, :file, path} -> Stream.resource( - fn -> open_excp_file(path) end, + fn -> open_excp_file(path, opts) end, &next_excp_item/1, &close_excp_file/1 ) @@ -378,15 +362,45 @@ defmodule ExCodecs.Spatial.Codec.Binary do end defp stream_from_binary(bin, opts) do - case decode(bin, opts) do - {:ok, %PointCloud{points: points}} -> - Stream.map(points, & &1) + if accel?(opts) do + stream_from_binary_accel(bin) + else + case decode(bin, Keyword.put(opts, :accel, false)) do + {:ok, %PointCloud{points: points}} -> Stream.map(points, & &1) + {:error, error} -> error_stream(error) + end + end + end - {:error, error} -> - error_stream(error) + defp stream_from_binary_accel( + <<@magic::binary, version::little-unsigned-16, flags::little-unsigned-16, + count::little-unsigned-64, body::binary>> + ) do + if version != @version do + error_stream( + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Unsupported binary point format version #{version}" + ) + ) + else + Stream.resource( + fn -> {:bin, body, flags, count, 0, 0} end, + &next_excp_accel/1, + fn _ -> :ok end + ) end end + defp stream_from_binary_accel(_) do + error_stream( + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Invalid ExCodecs binary point cloud" + ) + ) + end + defp resolve_source(bin, opts) do case Keyword.get(opts, :source, :auto) do :binary -> @@ -410,13 +424,93 @@ defmodule ExCodecs.Spatial.Codec.Binary do String.ends_with?(bin, [".excp", ".bin"])) end - defp open_excp_file(path) do + defp open_excp_file(path, opts) do + if accel?(opts) do + open_excp_mmap(path) + else + open_excp_io(path) + end + end + + defp open_excp_io(path) do case File.open(path, [:read, :binary, :raw]) do {:ok, io} -> parse_excp_header(io, IO.binread(io, 16)) {:error, reason} -> io_error(reason) end end + defp open_excp_mmap(path) do + with {:ok, header} <- read_file_prefix(path, 16), + {:ok, ref} <- Accel.mmap_open(path), + {:ok, state} <- mmap_state_from_header(ref, header) do + state + else + # coveralls-ignore-start + {:error, :nif_not_loaded} -> + open_excp_io(path) + + # coveralls-ignore-stop + {:error, reason} when is_atom(reason) -> + io_error(reason) + + {:error, %Error{}} = err -> + err + + # coveralls-ignore-start + {:error, other} -> + io_error(other) + # coveralls-ignore-stop + end + end + + defp mmap_state_from_header( + ref, + <<@magic::binary, version::little-unsigned-16, flags::little-unsigned-16, + count::little-unsigned-64>> + ) do + if version != @version do + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Unsupported binary point format version #{version}" + )} + else + {:ok, {:mmap, ref, flags, count, 16, 0}} + end + end + + defp mmap_state_from_header(_ref, _header) do + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Invalid ExCodecs binary point cloud" + )} + end + + defp read_file_prefix(path, n) do + case File.open(path, [:read, :binary, :raw]) do + {:ok, io} -> + data = IO.binread(io, n) + File.close(io) + + case data do + bin when is_binary(bin) -> + {:ok, bin} + + :eof -> + {:ok, <<>>} + + # coveralls-ignore-start + {:error, reason} -> + {:error, reason} + # coveralls-ignore-stop + end + + {:error, reason} -> + {:error, reason} + end + end + defp parse_excp_header( io, <<@magic::binary, version::little-unsigned-16, flags::little-unsigned-16, @@ -455,11 +549,14 @@ defmodule ExCodecs.Spatial.Codec.Binary do )} end + # coveralls-ignore-start defp parse_excp_header(io, {:error, reason}) do File.close(io) io_error(reason) end + # coveralls-ignore-stop + defp next_excp_item({:ok, io, _flags, count, _stride, i}) when i >= count do {:halt, {:done, io}} end @@ -471,8 +568,10 @@ defmodule ExCodecs.Spatial.Codec.Binary do {:ok, point, _} -> {[point], {:ok, io, flags, count, stride, i + 1}} + # coveralls-ignore-start {:error, error} -> {[{:error, error}], {:done, io}} + # coveralls-ignore-stop end :eof -> @@ -493,14 +592,108 @@ defmodule ExCodecs.Spatial.Codec.Binary do )} ], {:done, io}} + # coveralls-ignore-start {:error, reason} -> {[{:error, io_error_struct(reason)}], {:done, io}} + # coveralls-ignore-stop end end + defp next_excp_item({:mmap, _, _, count, _, i}) when i >= count, do: {:halt, :done_mmap} + + defp next_excp_item({:mmap, ref, flags, count, offset, i}) do + next_excp_accel({:mmap, ref, flags, count, offset, i}) + end + defp next_excp_item({:error, error}), do: {[{:error, error}], :done} defp next_excp_item({:done, _io}), do: {:halt, :done} defp next_excp_item(:done), do: {:halt, :done} + # coveralls-ignore-start + defp next_excp_item(:done_mmap), do: {:halt, :done} + + defp next_excp_accel(:done), do: {:halt, :done} + defp next_excp_accel(:done_mmap), do: {:halt, :done} + # coveralls-ignore-stop + + defp next_excp_accel({:bin, _body, _flags, count, _offset, i}) when i >= count do + {:halt, :done} + end + + defp next_excp_accel({:bin, body, flags, count, offset, i}) do + want = min(Accel.chunk_size(), count - i) + + case Accel.excp_unpack(body, flags, offset, want) do + {:ok, {[], _}} when want > 0 -> + {[ + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Truncated point record" + )} + ], :done} + + {:ok, {points, next}} -> + {points, {:bin, body, flags, count, next, i + length(points)}} + + # coveralls-ignore-start + {:error, :nif_not_loaded} -> + {[ + {:error, + Error.new(:nif_not_loaded, + codec: :spatial_binary, + message: "Spatial NIF unavailable mid-stream" + )} + ], :done} + + {:error, _} -> + {[ + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Truncated point record" + )} + ], :done} + + # coveralls-ignore-stop + end + end + + # coveralls-ignore-start + defp next_excp_accel({:mmap, _ref, _flags, count, _offset, i}) when i >= count do + {:halt, :done_mmap} + end + + # coveralls-ignore-stop + + defp next_excp_accel({:mmap, ref, flags, count, offset, i}) do + want = min(Accel.chunk_size(), count - i) + + case Accel.excp_unpack_mmap(ref, flags, offset, want) do + {:ok, {[], _}} when want > 0 -> + {[ + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Truncated point record" + )} + ], :done_mmap} + + {:ok, {points, next}} -> + {points, {:mmap, ref, flags, count, next, i + length(points)}} + + # coveralls-ignore-start + {:error, _} -> + {[ + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Truncated point record" + )} + ], :done_mmap} + + # coveralls-ignore-stop + end + end defp close_excp_file({:ok, io, _, _, _, _}), do: File.close(io) defp close_excp_file({:done, io}), do: File.close(io) @@ -585,9 +778,84 @@ defmodule ExCodecs.Spatial.Codec.Binary do end end - defp decode_points(bin, 0, _flags), do: {:ok, [], bin} + defp encode_points_body(points, flags, opts) do + if accel?(opts) do + case Accel.excp_pack(points, flags) do + {:ok, bin} -> + bin + + # coveralls-ignore-start + _ -> + encode_points_body_elixir(points, flags) + # coveralls-ignore-stop + end + else + encode_points_body_elixir(points, flags) + end + end + + defp encode_points_body_elixir(points, flags) do + IO.iodata_to_binary(Enum.map(points, &encode_point(&1, flags))) + end + + defp write_point_chunks(enumerable, io, flags, opts) do + chunk_size = if accel?(opts), do: Accel.chunk_size(), else: 1 + + enumerable + |> Stream.chunk_every(chunk_size) + |> Enum.reduce(0, fn chunk, n -> + points = assert_points!(chunk) + :ok = write_points_chunk(io, points, flags, opts) + n + length(points) + end) + end + + defp assert_points!(chunk) do + Enum.map(chunk, fn + %Point{} = p -> p + other -> throw({:bad_point, other}) + end) + end + + defp write_points_chunk(io, points, flags, opts) do + if accel?(opts) do + case Accel.excp_pack(points, flags) do + {:ok, bin} -> + IO.binwrite(io, bin) + + # coveralls-ignore-start + _ -> + Enum.each(points, &IO.binwrite(io, encode_point(&1, flags))) + # coveralls-ignore-stop + end + else + Enum.each(points, &IO.binwrite(io, encode_point(&1, flags))) + end + + :ok + end + + defp accel?(opts), do: Keyword.get(opts, :accel, true) != false and Accel.available?() + + defp decode_points(bin, 0, _flags, _opts), do: {:ok, [], bin} + + defp decode_points(bin, count, flags, opts) do + if accel?(opts) do + case Accel.excp_unpack(bin, flags, 0, count) do + {:ok, {points, _}} when length(points) == count -> + {:ok, points, <<>>} + + # Incomplete / failed Accel unpack: Elixir path preserves field-specific + # truncation messages (RGB / RGBA / normal). + _ -> + decode_points_elixir(bin, count, flags) + end + else + decode_points_elixir(bin, count, flags) + end + end - defp decode_points(bin, count, flags) do + defp decode_points_elixir(bin, count, flags) do Enum.reduce_while(1..count, {:ok, [], bin}, fn _, {:ok, acc, rest} -> case decode_point(rest, flags) do {:ok, point, next} -> {:cont, {:ok, [point | acc], next}} diff --git a/lib/ex_codecs/spatial/codec/gsplat.ex b/lib/ex_codecs/spatial/codec/gsplat.ex index 21d702d..f104383 100644 --- a/lib/ex_codecs/spatial/codec/gsplat.ex +++ b/lib/ex_codecs/spatial/codec/gsplat.ex @@ -22,6 +22,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do """ alias ExCodecs.Error + alias ExCodecs.Spatial.Accel alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Metadata} @magic "GSPL" @@ -76,7 +77,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do @spec encode(GaussianCloud.t(), keyword()) :: {:ok, binary()} | {:error, Error.t()} def encode(data, opts \\ []) - def encode(%GaussianCloud{gaussians: gaussians}, _opts) do + def encode(%GaussianCloud{gaussians: gaussians}, opts) do sh_rest = max_sh_rest(gaussians) flags = if sh_rest > 0, do: 1, else: 0 @@ -84,9 +85,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do <<@magic::binary, @version::little-unsigned-16, flags::little-unsigned-16, length(gaussians)::little-unsigned-64, sh_rest::little-unsigned-16>> - body = - IO.iodata_to_binary(Enum.map(gaussians, fn g -> encode_gaussian(g, sh_rest) end)) - + body = encode_gaussians_body(gaussians, sh_rest, opts) {:ok, header <> body} end @@ -143,15 +142,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do try do :ok = IO.binwrite(io, gspl_header(flags, 0, sh_rest)) - count = - Enum.reduce(enumerable, 0, fn - %Gaussian{} = g, n -> - :ok = IO.binwrite(io, encode_gaussian(g, sh_rest)) - n + 1 - - other, _n -> - throw({:bad_gaussian, other}) - end) + count = write_gaussian_chunks(enumerable, io, sh_rest, opts) {:ok, 0} = :file.position(io, 0) :ok = IO.binwrite(io, gspl_header(flags, count, sh_rest)) @@ -277,7 +268,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do def decode( <<@magic::binary, version::little-unsigned-16, _flags::little-unsigned-16, count::little-unsigned-64, sh_rest::little-unsigned-16, rest::binary>>, - _opts + opts ) do if version != @version do {:error, @@ -286,7 +277,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do message: "Unsupported GSPLAT version #{version}" )} else - with {:ok, gaussians, _} <- decode_gaussians(rest, count, sh_rest) do + with {:ok, gaussians, _} <- decode_gaussians(rest, count, sh_rest, opts) do meta = Metadata.new(entries: %{"format" => "gsplat", "version" => version}) {:ok, GaussianCloud.new(gaussians, metadata: meta)} end @@ -302,14 +293,14 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do ## Streaming behavior - * `source: :file` (or `:auto` when the argument looks like a path to a - regular file) reads the 18-byte header, then **one record at a time** - via `IO.binread/2`. Peak memory is O(header + one record). - * `source: :binary` (or `:auto` for an in-memory GSPL payload) still - materializes through `decode/2`, then yields the list. A future Rust - backend may memory-map large binaries instead. + * `source: :file` (or `:auto` path detection): header then chunked + records — Rust mmap + DirtyCpu unpack when available, otherwise + `IO.binread/2` per record. + * `source: :binary`: chunked Rust unpack when available; otherwise + materializes through `decode/2`. - Prefer `source: :file` for large `.gspl` paths. + Prefer `source: :file` for multi‑MB / multi‑GB `.gspl` paths. + Pass `accel: false` to force the pure-Elixir path. ## Arguments @@ -349,7 +340,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do {:ok, :file, path} -> Stream.resource( - fn -> open_gspl_file(path) end, + fn -> open_gspl_file(path, opts) end, &next_gspl_item/1, &close_gspl_file/1 ) @@ -357,15 +348,43 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do end defp stream_from_binary(bin, opts) do - case decode(bin, opts) do - {:ok, %GaussianCloud{gaussians: gs}} -> - Stream.map(gs, & &1) + if accel?(opts) do + stream_from_binary_accel(bin) + else + case decode(bin, Keyword.put(opts, :accel, false)) do + {:ok, %GaussianCloud{gaussians: gs}} -> + Stream.map(gs, & &1) - {:error, error} -> - error_stream(error) + {:error, error} -> + error_stream(error) + end end end + defp stream_from_binary_accel( + <<@magic::binary, version::little-unsigned-16, _flags::little-unsigned-16, + count::little-unsigned-64, sh_rest::little-unsigned-16, body::binary>> + ) do + if version != @version do + error_stream( + Error.new(:invalid_data, + codec: :gsplat, + message: "Unsupported GSPLAT version #{version}" + ) + ) + else + Stream.resource( + fn -> {:bin, body, sh_rest, count, 0, 0} end, + &next_gspl_accel/1, + fn _ -> :ok end + ) + end + end + + defp stream_from_binary_accel(_) do + error_stream(Error.new(:invalid_data, codec: :gsplat, message: "Invalid GSPLAT binary")) + end + defp resolve_source(bin, opts) do case Keyword.get(opts, :source, :auto) do :binary -> @@ -389,13 +408,89 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do String.ends_with?(bin, [".gspl", ".bin"])) end - defp open_gspl_file(path) do + defp open_gspl_file(path, opts) do + if accel?(opts) do + open_gspl_mmap(path) + else + open_gspl_io(path) + end + end + + defp open_gspl_io(path) do case File.open(path, [:read, :binary, :raw]) do {:ok, io} -> parse_gspl_header(io, IO.binread(io, 18)) {:error, reason} -> io_error(reason) end end + defp open_gspl_mmap(path) do + with {:ok, header} <- read_file_prefix(path, 18), + {:ok, ref} <- Accel.mmap_open(path), + {:ok, state} <- mmap_state_from_header(ref, header) do + state + else + # coveralls-ignore-start + {:error, :nif_not_loaded} -> + open_gspl_io(path) + + # coveralls-ignore-stop + {:error, reason} when is_atom(reason) -> + io_error(reason) + + {:error, %Error{}} = err -> + err + + # coveralls-ignore-start + {:error, other} -> + io_error(other) + # coveralls-ignore-stop + end + end + + defp mmap_state_from_header( + ref, + <<@magic::binary, version::little-unsigned-16, _flags::little-unsigned-16, + count::little-unsigned-64, sh_rest::little-unsigned-16>> + ) do + if version != @version do + {:error, + Error.new(:invalid_data, + codec: :gsplat, + message: "Unsupported GSPLAT version #{version}" + )} + else + {:ok, {:mmap, ref, sh_rest, count, 18, 0}} + end + end + + defp mmap_state_from_header(_ref, _header) do + {:error, Error.new(:invalid_data, codec: :gsplat, message: "Invalid GSPLAT binary")} + end + + defp read_file_prefix(path, n) do + case File.open(path, [:read, :binary, :raw]) do + {:ok, io} -> + data = IO.binread(io, n) + File.close(io) + + case data do + bin when is_binary(bin) -> + {:ok, bin} + + :eof -> + {:ok, <<>>} + + # coveralls-ignore-start + {:error, reason} -> + {:error, reason} + # coveralls-ignore-stop + end + + {:error, reason} -> + {:error, reason} + end + end + defp parse_gspl_header( io, <<@magic::binary, version::little-unsigned-16, _flags::little-unsigned-16, @@ -424,11 +519,14 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do {:error, Error.new(:invalid_data, codec: :gsplat, message: "Invalid GSPLAT binary")} end + # coveralls-ignore-start defp parse_gspl_header(io, {:error, reason}) do File.close(io) io_error(reason) end + # coveralls-ignore-stop + defp next_gspl_item({:ok, io, _sh_rest, count, _stride, i}) when i >= count do {:halt, {:done, io}} end @@ -440,8 +538,10 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do {:ok, gaussian, _} -> {[gaussian], {:ok, io, sh_rest, count, stride, i + 1}} + # coveralls-ignore-start {:error, error} -> {[{:error, error}], {:done, io}} + # coveralls-ignore-stop end :eof -> @@ -454,14 +554,108 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do {:error, Error.new(:invalid_data, codec: :gsplat, message: "Truncated Gaussian record")} ], {:done, io}} + # coveralls-ignore-start {:error, reason} -> {[{:error, io_error_struct(reason)}], {:done, io}} + # coveralls-ignore-stop end end + defp next_gspl_item({:mmap, _, _, count, _, i}) when i >= count, do: {:halt, :done_mmap} + + defp next_gspl_item({:mmap, ref, sh_rest, count, offset, i}) do + next_gspl_accel({:mmap, ref, sh_rest, count, offset, i}) + end + defp next_gspl_item({:error, error}), do: {[{:error, error}], :done} defp next_gspl_item({:done, _io}), do: {:halt, :done} defp next_gspl_item(:done), do: {:halt, :done} + # coveralls-ignore-start + defp next_gspl_item(:done_mmap), do: {:halt, :done} + + defp next_gspl_accel(:done), do: {:halt, :done} + defp next_gspl_accel(:done_mmap), do: {:halt, :done} + # coveralls-ignore-stop + + defp next_gspl_accel({:bin, _body, _sh_rest, count, _offset, i}) when i >= count do + {:halt, :done} + end + + defp next_gspl_accel({:bin, body, sh_rest, count, offset, i}) do + want = min(Accel.chunk_size(), count - i) + + case Accel.gspl_unpack(body, sh_rest, offset, want) do + {:ok, {[], _}} when want > 0 -> + {[ + {:error, + Error.new(:invalid_data, + codec: :gsplat, + message: "Truncated Gaussian record" + )} + ], :done} + + {:ok, {gaussians, next}} -> + {gaussians, {:bin, body, sh_rest, count, next, i + length(gaussians)}} + + # coveralls-ignore-start + {:error, :nif_not_loaded} -> + {[ + {:error, + Error.new(:nif_not_loaded, + codec: :gsplat, + message: "Spatial NIF unavailable mid-stream" + )} + ], :done} + + {:error, _} -> + {[ + {:error, + Error.new(:invalid_data, + codec: :gsplat, + message: "Truncated Gaussian record" + )} + ], :done} + + # coveralls-ignore-stop + end + end + + # coveralls-ignore-start + defp next_gspl_accel({:mmap, _ref, _sh_rest, count, _offset, i}) when i >= count do + {:halt, :done_mmap} + end + + # coveralls-ignore-stop + + defp next_gspl_accel({:mmap, ref, sh_rest, count, offset, i}) do + want = min(Accel.chunk_size(), count - i) + + case Accel.gspl_unpack_mmap(ref, sh_rest, offset, want) do + {:ok, {[], _}} when want > 0 -> + {[ + {:error, + Error.new(:invalid_data, + codec: :gsplat, + message: "Truncated Gaussian record" + )} + ], :done_mmap} + + {:ok, {gaussians, next}} -> + {gaussians, {:mmap, ref, sh_rest, count, next, i + length(gaussians)}} + + # coveralls-ignore-start + {:error, _} -> + {[ + {:error, + Error.new(:invalid_data, + codec: :gsplat, + message: "Truncated Gaussian record" + )} + ], :done_mmap} + + # coveralls-ignore-stop + end + end defp close_gspl_file({:ok, io, _, _, _, _}), do: File.close(io) defp close_gspl_file({:done, io}), do: File.close(io) @@ -528,11 +722,88 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do defp sh_rest_values(nil), do: [] defp sh_rest_values([_dc | rest]), do: List.flatten(rest) + # coveralls-ignore-start defp sh_rest_values(list) when is_list(list), do: list |> List.flatten() |> Enum.drop(3) + # coveralls-ignore-stop + + defp encode_gaussians_body(gaussians, sh_rest, opts) do + if accel?(opts) do + case Accel.gspl_pack(gaussians, sh_rest) do + {:ok, bin} -> + bin + + # coveralls-ignore-start + _ -> + encode_gaussians_body_elixir(gaussians, sh_rest) + # coveralls-ignore-stop + end + else + encode_gaussians_body_elixir(gaussians, sh_rest) + end + end + + defp encode_gaussians_body_elixir(gaussians, sh_rest) do + IO.iodata_to_binary(Enum.map(gaussians, fn g -> encode_gaussian(g, sh_rest) end)) + end - defp decode_gaussians(bin, 0, _sh_rest), do: {:ok, [], bin} + defp write_gaussian_chunks(enumerable, io, sh_rest, opts) do + chunk_size = if accel?(opts), do: Accel.chunk_size(), else: 1 + + enumerable + |> Stream.chunk_every(chunk_size) + |> Enum.reduce(0, fn chunk, n -> + gaussians = assert_gaussians!(chunk) + :ok = write_gaussian_chunk(io, gaussians, sh_rest, opts) + n + length(gaussians) + end) + end + + defp assert_gaussians!(chunk) do + Enum.map(chunk, fn + %Gaussian{} = g -> g + other -> throw({:bad_gaussian, other}) + end) + end + + defp write_gaussian_chunk(io, gaussians, sh_rest, opts) do + if accel?(opts) do + case Accel.gspl_pack(gaussians, sh_rest) do + {:ok, bin} -> + IO.binwrite(io, bin) + + # coveralls-ignore-start + _ -> + Enum.each(gaussians, &IO.binwrite(io, encode_gaussian(&1, sh_rest))) + # coveralls-ignore-stop + end + else + Enum.each(gaussians, &IO.binwrite(io, encode_gaussian(&1, sh_rest))) + end + + :ok + end + + defp accel?(opts), do: Keyword.get(opts, :accel, true) != false and Accel.available?() + + defp decode_gaussians(bin, 0, _sh_rest, _opts), do: {:ok, [], bin} + + defp decode_gaussians(bin, count, sh_rest, opts) do + if accel?(opts) do + case Accel.gspl_unpack(bin, sh_rest, 0, count) do + {:ok, {gaussians, _}} when length(gaussians) == count -> + {:ok, gaussians, <<>>} + + # Incomplete / failed Accel unpack: Elixir path preserves "Truncated SH" + # and other field-specific messages. + _ -> + decode_gaussians_elixir(bin, count, sh_rest) + end + else + decode_gaussians_elixir(bin, count, sh_rest) + end + end - defp decode_gaussians(bin, count, sh_rest) do + defp decode_gaussians_elixir(bin, count, sh_rest) do Enum.reduce_while(1..count, {:ok, [], bin}, fn _, {:ok, acc, rest} -> case decode_gaussian(rest, sh_rest) do {:ok, g, next} -> {:cont, {:ok, [g | acc], next}} diff --git a/lib/ex_codecs/spatial/codec/ply.ex b/lib/ex_codecs/spatial/codec/ply.ex index 726c4dd..eebac3f 100644 --- a/lib/ex_codecs/spatial/codec/ply.ex +++ b/lib/ex_codecs/spatial/codec/ply.ex @@ -14,6 +14,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do """ alias ExCodecs.Error + alias ExCodecs.Spatial.Accel alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Metadata, Point, PointCloud} @typedoc """ @@ -184,7 +185,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do with {:ok, header, body} <- split_header(data), {:ok, parsed} <- parse_header(header) do - decode_parsed_body(resolve_as(as, parsed.properties), body, parsed) + decode_parsed_body(resolve_as(as, parsed.properties), body, parsed, opts) end end @@ -712,14 +713,14 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end - defp decode_parsed_body(:point_cloud, body, parsed) do - with {:ok, points} <- decode_vertices(body, parsed) do + defp decode_parsed_body(:point_cloud, body, parsed, opts) do + with {:ok, points} <- decode_vertices(body, parsed, opts) do {:ok, PointCloud.new(points, metadata: ply_metadata(parsed))} end end - defp decode_parsed_body(:gaussian_cloud, body, parsed) do - with {:ok, gaussians} <- decode_gaussians(body, parsed) do + defp decode_parsed_body(:gaussian_cloud, body, parsed, opts) do + with {:ok, gaussians} <- decode_gaussians(body, parsed, opts) do {:ok, GaussianCloud.new(gaussians, metadata: ply_metadata(parsed))} end end @@ -731,7 +732,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do ) end - defp decode_vertices(body, %{format: :ascii, count: count, properties: props}) do + defp decode_vertices(body, %{format: :ascii, count: count, properties: props}, _opts) do lines = body |> String.split(~r/\r\n|\n|\r/, trim: true) @@ -754,7 +755,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end - defp decode_vertices(body, %{format: endian, count: count, properties: props}) + defp decode_vertices(body, %{format: endian, count: count, properties: props}, opts) when endian in [:binary_le, :binary_be] do stride = Enum.reduce(props, 0, fn p, acc -> acc + type_size(p.type) end) expected = stride * count @@ -766,18 +767,46 @@ defmodule ExCodecs.Spatial.Codec.PLY do message: "Binary PLY body too short" )} else - {points, _} = - Enum.map_reduce(1..count, body, fn _, rest -> - {values, next} = unpack_row(rest, props, endian) - {values_to_point(values, props), next} - end) + decode_vertices_binary(body, props, endian, count, expected, opts) + end + end - {:ok, points} + defp decode_vertices_binary(body, props, endian, count, expected, opts) do + if accel?(opts) do + types = Enum.map(props, & &1.type) + + case Accel.ply_binary_unpack(body, types, endian, 0, count) do + {:ok, {rows, next}} when length(rows) == count and next >= expected -> + {:ok, Enum.map(rows, &values_to_point(&1, props))} + + # coveralls-ignore-start + {:ok, _} -> + decode_vertices_binary_loop(body, props, endian, count) + + {:error, :nif_not_loaded} -> + decode_vertices_binary_loop(body, props, endian, count) + + {:error, _} -> + decode_vertices_binary_loop(body, props, endian, count) + # coveralls-ignore-stop + end + else + decode_vertices_binary_loop(body, props, endian, count) end end - defp decode_gaussians(body, parsed) do - with {:ok, points} <- decode_vertices(body, parsed) do + defp decode_vertices_binary_loop(body, props, endian, count) do + {points, _} = + Enum.map_reduce(1..count, body, fn _, rest -> + {values, next} = unpack_row(rest, props, endian) + {values_to_point(values, props), next} + end) + + {:ok, points} + end + + defp decode_gaussians(body, parsed, opts) do + with {:ok, points} <- decode_vertices(body, parsed, opts) do {:ok, Enum.map(points, &point_to_gaussian/1)} end end @@ -941,10 +970,10 @@ defmodule ExCodecs.Spatial.Codec.PLY do defp open_ply_stream(path, opts) do case File.open(path, [:read, :binary, :raw]) do {:ok, io} -> - case scan_ply_header(io, "") do - {:ok, parsed, leftover} -> + case scan_ply_header(io, "", 0) do + {:ok, parsed, leftover, body_offset} -> as = resolve_as(Keyword.get(opts, :as, :auto), parsed.properties) - {:ok, init_ply_stream_state(io, leftover, parsed, as)} + {:ok, init_ply_stream_state(path, io, leftover, parsed, as, opts, body_offset)} {:error, error} -> File.close(io) @@ -956,38 +985,72 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end - defp init_ply_stream_state(io, leftover, %{format: :ascii} = parsed, as) do + defp init_ply_stream_state( + _path, + io, + leftover, + %{format: :ascii} = parsed, + as, + _opts, + _body_offset + ) do {:ascii, io, leftover, parsed, as, 0} end - defp init_ply_stream_state(io, leftover, %{format: endian} = parsed, as) + defp init_ply_stream_state(path, io, leftover, %{format: endian} = parsed, as, opts, body_offset) when endian in [:binary_le, :binary_be] do + if accel?(opts) do + case Accel.mmap_open(path) do + {:ok, ref} -> + File.close(io) + types = Enum.map(parsed.properties, & &1.type) + {:binary_mmap, ref, types, endian, parsed, as, body_offset, 0, 0} + + # coveralls-ignore-start + {:error, :nif_not_loaded} -> + init_ply_binary_io_stream(io, leftover, parsed, as, endian) + + {:error, _} -> + init_ply_binary_io_stream(io, leftover, parsed, as, endian) + # coveralls-ignore-stop + end + else + init_ply_binary_io_stream(io, leftover, parsed, as, endian) + end + end + + defp init_ply_binary_io_stream(io, leftover, parsed, as, endian) do stride = Enum.reduce(parsed.properties, 0, fn p, acc -> acc + type_size(p.type) end) {:binary, io, leftover, endian, parsed, as, stride, 0} end - defp scan_ply_header(io, acc) do + defp scan_ply_header(io, acc, bytes_read) do case IO.binread(io, @header_scan_chunk) do data when is_binary(data) and byte_size(data) > 0 -> buf = acc <> data - match_or_continue_header(io, buf) + match_or_continue_header(io, buf, bytes_read + byte_size(data)) + # coveralls-ignore-start {:error, reason} -> ply_io_error(reason) + # coveralls-ignore-stop _eof_or_empty -> header_eof_error(acc) end end - defp match_or_continue_header(io, buf) do + defp match_or_continue_header(io, buf, bytes_read) do case :binary.match(buf, "end_header") do {pos, len} -> + buf_start = bytes_read - byte_size(buf) header = binary_part(buf, 0, pos + len) rest = binary_part(buf, pos + len, byte_size(buf) - pos - len) + {body_rest, skip} = strip_leading_newlines_count(rest) + body_offset = buf_start + pos + len + skip with {:ok, parsed} <- parse_header(header) do - {:ok, parsed, strip_leading_newlines(rest)} + {:ok, parsed, body_rest, body_offset} end :nomatch when byte_size(buf) > @header_max_bytes -> @@ -998,16 +1061,69 @@ defmodule ExCodecs.Spatial.Codec.PLY do )} :nomatch -> - scan_ply_header(io, buf) + scan_ply_header(io, buf, bytes_read) end end + defp strip_leading_newlines_count(<<"\r\n", rest::binary>>), do: {rest, 2} + defp strip_leading_newlines_count(<<"\n", rest::binary>>), do: {rest, 1} + defp strip_leading_newlines_count(<<"\r", rest::binary>>), do: {rest, 1} + defp strip_leading_newlines_count(bin), do: {bin, 0} + defp header_eof_error(_acc) do # Acc never contains a complete end_header here: that case is handled in # match_or_continue_header/2 while chunks are still arriving. {:error, Error.new(:invalid_data, codec: :ply, message: "PLY header missing end_header")} end + defp next_ply_stream_item( + {:ok, {:binary_mmap, _ref, _types, _endian, parsed, _as, _body_offset, _offset, i}} + ) + when i >= parsed.count do + {:halt, :done_mmap} + end + + defp next_ply_stream_item( + {:ok, {:binary_mmap, ref, types, endian, parsed, as, body_offset, offset, i}} + ) do + want = min(Accel.chunk_size(), parsed.count - i) + + case Accel.ply_binary_unpack_mmap(ref, types, endian, body_offset + offset, want) do + {:ok, {[], _next}} when want > 0 -> + {[ + {:error, + Error.new(:invalid_data, + codec: :ply, + message: "Binary PLY body too short" + )} + ], :done_mmap} + + {:ok, {rows, next_offset}} -> + items = + Enum.map(rows, fn values -> + emit_vertex(values_to_point(values, parsed.properties), as) + end) + + new_offset = next_offset - body_offset + + {items, + {:ok, + {:binary_mmap, ref, types, endian, parsed, as, body_offset, new_offset, i + length(rows)}}} + + # coveralls-ignore-start + {:error, _} -> + {[ + {:error, + Error.new(:invalid_data, + codec: :ply, + message: "Binary PLY body too short" + )} + ], :done_mmap} + + # coveralls-ignore-stop + end + end + defp next_ply_stream_item({:ok, {:binary, io, _buf, _endian, parsed, _as, _stride, i}}) when i >= parsed.count do {:halt, {:done, io}} @@ -1044,12 +1160,17 @@ defmodule ExCodecs.Spatial.Codec.PLY do defp next_ply_stream_item({:error, error}), do: {[{:error, error}], :done} defp next_ply_stream_item({:done, _io}), do: {:halt, :done} defp next_ply_stream_item(:done), do: {:halt, :done} + defp next_ply_stream_item(:done_mmap), do: {:halt, :done_mmap} defp close_ply_stream({:ok, {:binary, io, _, _, _, _, _, _}}), do: File.close(io) defp close_ply_stream({:ok, {:ascii, io, _, _, _, _}}), do: File.close(io) + defp close_ply_stream({:ok, {:binary_mmap, _, _, _, _, _, _, _, _}}), do: :ok defp close_ply_stream({:done, io}), do: File.close(io) + defp close_ply_stream(:done_mmap), do: :ok defp close_ply_stream(_), do: :ok + defp accel?(opts), do: Keyword.get(opts, :accel, true) != false and Accel.available?() + defp emit_vertex(point, :point_cloud), do: point defp emit_vertex(point, :gaussian_cloud), do: point_to_gaussian(point) @@ -1063,9 +1184,11 @@ defmodule ExCodecs.Spatial.Codec.PLY do data when is_binary(data) and byte_size(data) > 0 -> take_exact_bytes(io, buf <> data, n) + # coveralls-ignore-start {:error, reason} -> ply_io_error(reason) + # coveralls-ignore-stop _eof_or_short -> {:error, Error.new(:invalid_data, codec: :ply, message: "Binary PLY body too short")} end @@ -1090,9 +1213,11 @@ defmodule ExCodecs.Spatial.Codec.PLY do data when is_binary(data) and byte_size(data) > 0 -> take_ascii_vertex_line(io, buf <> data) + # coveralls-ignore-start {:error, reason} -> ply_io_error(reason) + # coveralls-ignore-stop _eof_or_empty -> trimmed = String.trim(buf) diff --git a/livebooks/01_introduction.livemd b/livebooks/01_introduction.livemd index 9e3017f..4c9f23d 100644 --- a/livebooks/01_introduction.livemd +++ b/livebooks/01_introduction.livemd @@ -13,7 +13,7 @@ # ) Mix.install([ - {:ex_codecs, "~> 0.2"} + {:ex_codecs, "~> 0.2.3"} ]) ``` diff --git a/livebooks/02_compression_fundamentals.livemd b/livebooks/02_compression_fundamentals.livemd index ba07d8b..34c01a0 100644 --- a/livebooks/02_compression_fundamentals.livemd +++ b/livebooks/02_compression_fundamentals.livemd @@ -16,7 +16,7 @@ # ) Mix.install( [ - {:ex_codecs, "~> 0.2.0"}, + {:ex_codecs, "~> 0.2.3"}, {:jason, "~> 1.4"}, {:kino, "~> 0.14"}, {:kino_vega_lite, "~> 0.1.13"} diff --git a/livebooks/03_codec_comparison.livemd b/livebooks/03_codec_comparison.livemd index e2b1edc..67a5f48 100644 --- a/livebooks/03_codec_comparison.livemd +++ b/livebooks/03_codec_comparison.livemd @@ -16,7 +16,7 @@ # ) Mix.install( [ - {:ex_codecs, "~> 0.2"}, + {:ex_codecs, "~> 0.2.3"}, {:jason, "~> 1.4"}, {:kino, "~> 0.14"}, {:kino_vega_lite, "~> 0.1.13"} diff --git a/livebooks/04_building_storage_systems.livemd b/livebooks/04_building_storage_systems.livemd index 86d0cdc..81ac5a8 100644 --- a/livebooks/04_building_storage_systems.livemd +++ b/livebooks/04_building_storage_systems.livemd @@ -16,7 +16,7 @@ # ) Mix.install( [ - {:ex_codecs, "~> 0.2"}, + {:ex_codecs, "~> 0.2.3"}, {:jason, "~> 1.4"}, {:kino, "~> 0.14"}, {:kino_vega_lite, "~> 0.1.13"} diff --git a/livebooks/05_zarr_style_workloads.livemd b/livebooks/05_zarr_style_workloads.livemd index d126f31..e7920ca 100644 --- a/livebooks/05_zarr_style_workloads.livemd +++ b/livebooks/05_zarr_style_workloads.livemd @@ -14,7 +14,7 @@ # ) Mix.install( [ - {:ex_codecs, "~> 0.2"}, + {:ex_codecs, "~> 0.2.3"}, {:jason, "~> 1.4"}, {:kino, "~> 0.14"}, {:kino_vega_lite, "~> 0.1.13"} diff --git a/livebooks/06_spatial_codecs.livemd b/livebooks/06_spatial_codecs.livemd index 1c5dd49..4a61d8b 100644 --- a/livebooks/06_spatial_codecs.livemd +++ b/livebooks/06_spatial_codecs.livemd @@ -11,7 +11,7 @@ # ) Mix.install([ - {:ex_codecs, "~> 0.2"} + {:ex_codecs, "~> 0.2.3"} ]) ``` @@ -181,9 +181,10 @@ gaussian_cloud = ## Enumerable Helpers -The current `stream_encode/2` and `stream_decode/2` names describe an -enumerable-facing API. In v0.2.0 they still collect the complete enumerable or -payload in memory; they are not incremental file I/O. +`stream_decode/2` with `source: :file` reads one record/vertex at a time +(bounded memory; mmap + Rust unpack when available). `stream_encode/2` still +collects the enumerable, then encodes once — use `encode_to_file/3` with an +explicit `:schema` for EXCP/GSPL incremental writes. ```elixir {:ok, streamed_payload} = diff --git a/mix.exs b/mix.exs index 88ce789..9ae0f1f 100644 --- a/mix.exs +++ b/mix.exs @@ -1,7 +1,7 @@ defmodule ExCodecs.MixProject do use Mix.Project - @version "0.2.2" + @version "0.2.3" @source_url "https://github.com/thanos/codecs" def project do @@ -20,8 +20,13 @@ defmodule ExCodecs.MixProject do tool: ExCoveralls, ignore_modules: [ExCodecs.Native], threshold: 95 - ], - preferred_cli_env: [ + ] + ] + end + + def cli do + [ + preferred_envs: [ coveralls: :test, "coveralls.detail": :test, "coveralls.github": :test, diff --git a/native/ex_codecs_native/Cargo.lock b/native/ex_codecs_native/Cargo.lock index cba4fec..7fb4f41 100644 --- a/native/ex_codecs_native/Cargo.lock +++ b/native/ex_codecs_native/Cargo.lock @@ -77,12 +77,13 @@ checksum = "91622ff5e7162018101f2fea40d6ebf4a78bbe5a49736a2020649edf9693679e" [[package]] name = "ex_codecs_native" -version = "0.2.0" +version = "0.2.3" dependencies = [ "blosc2-pure-rs", "bzip2", "flate2", "lz4_flex", + "memmap2", "rustler", "snap", "structured-zstd", @@ -154,6 +155,15 @@ dependencies = [ "twox-hash", ] +[[package]] +name = "memmap2" +version = "0.9.11" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "d1219ed1b7f229ee7104d281dd01d6802fe28bb6e95d292942c4daacdeb798c0" +dependencies = [ + "libc", +] + [[package]] name = "miniz_oxide" version = "0.8.9" diff --git a/native/ex_codecs_native/Cargo.toml b/native/ex_codecs_native/Cargo.toml index 48a44f1..68a1c98 100644 --- a/native/ex_codecs_native/Cargo.toml +++ b/native/ex_codecs_native/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "ex_codecs_native" -version = "0.2.0" +version = "0.2.3" authors = ["ExCodecs Team"] edition = "2021" # structured-zstd declares MSRV 1.92 but calls the now-safe __cpuid without @@ -23,6 +23,7 @@ flate2 = { version = "1", default-features = false, features = ["rust_backend"] bzip2 = "0.6" # Official Blosc2 chunk format (BloscLZ, LZ4, LZ4HC, Zlib, Zstd + filters). blosc2-pure-rs = { version = "0.2.4", default-features = false, features = ["zlib-rs"] } +memmap2 = "0.9" [features] default = ["nif_version_2_17"] diff --git a/native/ex_codecs_native/src/lib.rs b/native/ex_codecs_native/src/lib.rs index a1e6118..e1460f9 100644 --- a/native/ex_codecs_native/src/lib.rs +++ b/native/ex_codecs_native/src/lib.rs @@ -3,6 +3,7 @@ mod blosc2_codec; mod bzip2_codec; mod lz4_codec; mod snappy_codec; +mod spatial; mod util; mod zstd_codec; @@ -16,5 +17,6 @@ fn codec_versions() -> std::collections::HashMap<&'static str, String> { versions.insert("snappy", snappy_codec::version()); versions.insert("bzip2", bzip2_codec::version()); versions.insert("blosc2", blosc2_codec::version()); + versions.insert("spatial", "excp-gspl-ply-1".to_string()); versions } diff --git a/native/ex_codecs_native/src/spatial.rs b/native/ex_codecs_native/src/spatial.rs new file mode 100644 index 0000000..539b057 --- /dev/null +++ b/native/ex_codecs_native/src/spatial.rs @@ -0,0 +1,621 @@ +//! Spatial binary hot paths: EXCP / GSPL / binary PLY unpack & pack. +//! +//! Wire formats match `docs/spatial_formats.md` and the pure-Elixir codecs. + +use std::fs::File; +use std::io::Write; + +use memmap2::Mmap; +use rustler::types::atom; +use rustler::{Binary, Encoder, Env, NifResult, Resource, ResourceArc, Term}; + +use crate::atoms; +use crate::util::{err, ok_binary}; + +const FLAG_COLOR: u16 = 0b001; +const FLAG_ALPHA: u16 = 0b010; +const FLAG_NORMAL: u16 = 0b100; + +const PLY_CHAR: u8 = 1; +const PLY_UCHAR: u8 = 2; +const PLY_SHORT: u8 = 3; +const PLY_USHORT: u8 = 4; +const PLY_INT: u8 = 5; +const PLY_UINT: u8 = 6; +const PLY_FLOAT: u8 = 7; +const PLY_DOUBLE: u8 = 8; + +pub struct MappedSpatial { + mmap: Mmap, +} + +#[rustler::resource_impl] +impl Resource for MappedSpatial {} + +fn nil_term(env: Env) -> Term { + atom::nil().encode(env) +} + +fn term_is_nil(term: Term) -> bool { + term.decode::() + .map(|a| a == atom::nil()) + .unwrap_or(false) +} + +fn read_f32_le(data: &[u8], off: usize) -> Option { + data.get(off..off + 4) + .and_then(|b| b.try_into().ok()) + .map(f32::from_le_bytes) +} + +fn read_f32_be(data: &[u8], off: usize) -> Option { + data.get(off..off + 4) + .and_then(|b| b.try_into().ok()) + .map(f32::from_be_bytes) +} + +fn read_f64_le(data: &[u8], off: usize) -> Option { + data.get(off..off + 8) + .and_then(|b| b.try_into().ok()) + .map(f64::from_le_bytes) +} + +fn read_f64_be(data: &[u8], off: usize) -> Option { + data.get(off..off + 8) + .and_then(|b| b.try_into().ok()) + .map(f64::from_be_bytes) +} + +fn read_i16_le(data: &[u8], off: usize) -> Option { + data.get(off..off + 2) + .and_then(|b| b.try_into().ok()) + .map(i16::from_le_bytes) +} + +fn read_i16_be(data: &[u8], off: usize) -> Option { + data.get(off..off + 2) + .and_then(|b| b.try_into().ok()) + .map(i16::from_be_bytes) +} + +fn read_u16_le(data: &[u8], off: usize) -> Option { + data.get(off..off + 2) + .and_then(|b| b.try_into().ok()) + .map(u16::from_le_bytes) +} + +fn read_u16_be(data: &[u8], off: usize) -> Option { + data.get(off..off + 2) + .and_then(|b| b.try_into().ok()) + .map(u16::from_be_bytes) +} + +fn read_i32_le(data: &[u8], off: usize) -> Option { + data.get(off..off + 4) + .and_then(|b| b.try_into().ok()) + .map(i32::from_le_bytes) +} + +fn read_i32_be(data: &[u8], off: usize) -> Option { + data.get(off..off + 4) + .and_then(|b| b.try_into().ok()) + .map(i32::from_be_bytes) +} + +fn read_u32_le(data: &[u8], off: usize) -> Option { + data.get(off..off + 4) + .and_then(|b| b.try_into().ok()) + .map(u32::from_le_bytes) +} + +fn read_u32_be(data: &[u8], off: usize) -> Option { + data.get(off..off + 4) + .and_then(|b| b.try_into().ok()) + .map(u32::from_be_bytes) +} + +fn excp_stride(flags: u16) -> usize { + let color = if flags & FLAG_ALPHA != 0 { + 4 + } else if flags & FLAG_COLOR != 0 { + 3 + } else { + 0 + }; + let normal = if flags & FLAG_NORMAL != 0 { 12 } else { 0 }; + 12 + color + normal +} + +fn gspl_stride(sh_rest: u16) -> usize { + 56 + (sh_rest as usize) * 4 +} + +fn ply_type_size(t: u8) -> Option { + match t { + PLY_CHAR | PLY_UCHAR => Some(1), + PLY_SHORT | PLY_USHORT => Some(2), + PLY_INT | PLY_UINT | PLY_FLOAT => Some(4), + PLY_DOUBLE => Some(8), + _ => None, + } +} + +fn ply_stride(types: &[u8]) -> Option { + let mut n = 0usize; + for t in types { + n = n.checked_add(ply_type_size(*t)?)?; + } + Some(n) +} + +fn ok_chunk<'a>(env: Env<'a>, records: Vec>, next_offset: u64) -> Term<'a> { + (atoms::ok(), (records, next_offset)).encode(env) +} + +fn decode_excp_point<'a>( + env: Env<'a>, + data: &[u8], + mut off: usize, + flags: u16, +) -> Option<(Term<'a>, usize)> { + let x = read_f32_le(data, off)? as f64; + let y = read_f32_le(data, off + 4)? as f64; + let z = read_f32_le(data, off + 8)? as f64; + off += 12; + + let color = if flags & FLAG_ALPHA != 0 { + let r = *data.get(off)? as i64; + let g = *data.get(off + 1)? as i64; + let b = *data.get(off + 2)? as i64; + let a = *data.get(off + 3)? as i64; + off += 4; + (r, g, b, a).encode(env) + } else if flags & FLAG_COLOR != 0 { + let r = *data.get(off)? as i64; + let g = *data.get(off + 1)? as i64; + let b = *data.get(off + 2)? as i64; + off += 3; + (r, g, b).encode(env) + } else { + nil_term(env) + }; + + let normal = if flags & FLAG_NORMAL != 0 { + let nx = read_f32_le(data, off)? as f64; + let ny = read_f32_le(data, off + 4)? as f64; + let nz = read_f32_le(data, off + 8)? as f64; + off += 12; + (nx, ny, nz).encode(env) + } else { + nil_term(env) + }; + + Some(((x, y, z, color, normal).encode(env), off)) +} + +fn unpack_excp_slice<'a>( + env: Env<'a>, + data: &[u8], + flags: u16, + offset: usize, + max_count: usize, +) -> Term<'a> { + let stride = excp_stride(flags); + let mut off = offset; + let mut records = Vec::with_capacity(max_count.min(8192)); + + for _ in 0..max_count { + if off.saturating_add(stride) > data.len() { + break; + } + match decode_excp_point(env, data, off, flags) { + Some((term, next)) => { + records.push(term); + off = next; + } + None => return err(env, atoms::invalid_data()), + } + } + + ok_chunk(env, records, off as u64) +} + +fn decode_gspl_point<'a>( + env: Env<'a>, + data: &[u8], + mut off: usize, + sh_rest: u16, +) -> Option<(Term<'a>, usize)> { + let mut vals = [0.0f64; 14]; + for v in &mut vals { + *v = read_f32_le(data, off)? as f64; + off += 4; + } + + let mut sh = Vec::with_capacity(sh_rest as usize); + for _ in 0..sh_rest { + sh.push(read_f32_le(data, off)? as f64); + off += 4; + } + + // Nested tuples stay within Rustler's Encoder arity limits. + let term = ( + (vals[0], vals[1], vals[2]), + (vals[3], vals[4], vals[5]), + vals[6], + (vals[7], vals[8], vals[9]), + (vals[10], vals[11], vals[12], vals[13]), + sh, + ) + .encode(env); + Some((term, off)) +} + +fn unpack_gspl_slice<'a>( + env: Env<'a>, + data: &[u8], + sh_rest: u16, + offset: usize, + max_count: usize, +) -> Term<'a> { + let stride = gspl_stride(sh_rest); + let mut off = offset; + let mut records = Vec::with_capacity(max_count.min(8192)); + + for _ in 0..max_count { + if off.saturating_add(stride) > data.len() { + break; + } + match decode_gspl_point(env, data, off, sh_rest) { + Some((term, next)) => { + records.push(term); + off = next; + } + None => return err(env, atoms::invalid_data()), + } + } + + ok_chunk(env, records, off as u64) +} + +fn decode_ply_value<'a>( + env: Env<'a>, + data: &[u8], + off: usize, + ty: u8, + little: bool, +) -> Option<(Term<'a>, usize)> { + match ty { + PLY_CHAR => { + let v = *data.get(off)? as i8 as i64; + Some((v.encode(env), off + 1)) + } + PLY_UCHAR => { + let v = *data.get(off)? as i64; + Some((v.encode(env), off + 1)) + } + PLY_SHORT => { + let v = if little { + read_i16_le(data, off)? + } else { + read_i16_be(data, off)? + } as i64; + Some((v.encode(env), off + 2)) + } + PLY_USHORT => { + let v = if little { + read_u16_le(data, off)? + } else { + read_u16_be(data, off)? + } as i64; + Some((v.encode(env), off + 2)) + } + PLY_INT => { + let v = if little { + read_i32_le(data, off)? + } else { + read_i32_be(data, off)? + } as i64; + Some((v.encode(env), off + 4)) + } + PLY_UINT => { + let v = if little { + read_u32_le(data, off)? + } else { + read_u32_be(data, off)? + } as i64; + Some((v.encode(env), off + 4)) + } + PLY_FLOAT => { + let v = if little { + read_f32_le(data, off)? + } else { + read_f32_be(data, off)? + } as f64; + Some((v.encode(env), off + 4)) + } + PLY_DOUBLE => { + let v = if little { + read_f64_le(data, off)? + } else { + read_f64_be(data, off)? + }; + Some((v.encode(env), off + 8)) + } + _ => None, + } +} + +fn unpack_ply_slice<'a>( + env: Env<'a>, + data: &[u8], + types: &[u8], + little: bool, + offset: usize, + max_count: usize, +) -> Term<'a> { + let Some(stride) = ply_stride(types) else { + return err(env, atoms::invalid_options()); + }; + + let mut off = offset; + let mut records = Vec::with_capacity(max_count.min(8192)); + + for _ in 0..max_count { + if off.saturating_add(stride) > data.len() { + break; + } + let start = off; + let mut row = Vec::with_capacity(types.len()); + for ty in types { + match decode_ply_value(env, data, off, *ty, little) { + Some((term, next)) => { + row.push(term); + off = next; + } + None => return err(env, atoms::invalid_data()), + } + } + if off - start != stride { + return err(env, atoms::invalid_data()); + } + records.push(row.encode(env)); + } + + ok_chunk(env, records, off as u64) +} + +fn write_f32_le(buf: &mut Vec, v: f64) { + buf.extend_from_slice(&(v as f32).to_le_bytes()); +} + +fn clamp_u8(v: i64) -> u8 { + v.clamp(0, 255) as u8 +} + +/// Point term: `{x, y, z, color, normal}` with color/normal nil or tuples. +fn pack_excp_point(buf: &mut Vec, term: Term, flags: u16) -> Result<(), ()> { + let (x, y, z, color, normal): (f64, f64, f64, Term, Term) = term.decode().map_err(|_| ())?; + write_f32_le(buf, x); + write_f32_le(buf, y); + write_f32_le(buf, z); + + if flags & FLAG_ALPHA != 0 { + let (r, g, b, a) = if term_is_nil(color) { + (0i64, 0, 0, 255) + } else if let Ok((r, g, b, a)) = color.decode::<(i64, i64, i64, i64)>() { + (r, g, b, a) + } else if let Ok((r, g, b)) = color.decode::<(i64, i64, i64)>() { + (r, g, b, 255) + } else { + return Err(()); + }; + buf.extend_from_slice(&[clamp_u8(r), clamp_u8(g), clamp_u8(b), clamp_u8(a)]); + } else if flags & FLAG_COLOR != 0 { + let (r, g, b) = if term_is_nil(color) { + (0i64, 0, 0) + } else if let Ok((r, g, b, _)) = color.decode::<(i64, i64, i64, i64)>() { + (r, g, b) + } else if let Ok((r, g, b)) = color.decode::<(i64, i64, i64)>() { + (r, g, b) + } else { + return Err(()); + }; + buf.extend_from_slice(&[clamp_u8(r), clamp_u8(g), clamp_u8(b)]); + } + + if flags & FLAG_NORMAL != 0 { + let (nx, ny, nz) = if term_is_nil(normal) { + (0.0, 0.0, 0.0) + } else { + normal.decode::<(f64, f64, f64)>().map_err(|_| ())? + }; + write_f32_le(buf, nx); + write_f32_le(buf, ny); + write_f32_le(buf, nz); + } + + Ok(()) +} + +fn pack_gspl_point(buf: &mut Vec, term: Term, sh_rest: u16) -> Result<(), ()> { + let (pos, color, opacity, scale, rot, sh): ( + (f64, f64, f64), + (f64, f64, f64), + f64, + (f64, f64, f64), + (f64, f64, f64, f64), + Vec, + ) = term.decode().map_err(|_| ())?; + + let (x, y, z) = pos; + let (r, g, b) = color; + let (sx, sy, sz) = scale; + let (rw, rx, ry, rz) = rot; + + for v in [x, y, z, r, g, b, opacity, sx, sy, sz, rw, rx, ry, rz] { + write_f32_le(buf, v); + } + + for i in 0..sh_rest as usize { + write_f32_le(buf, sh.get(i).copied().unwrap_or(0.0)); + } + + Ok(()) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn excp_unpack<'a>( + env: Env<'a>, + data: Binary, + flags: u16, + offset: u64, + max_count: u64, +) -> Term<'a> { + unpack_excp_slice( + env, + data.as_slice(), + flags, + offset as usize, + max_count as usize, + ) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn excp_pack<'a>(env: Env<'a>, records: Vec, flags: u16) -> Term<'a> { + let mut buf = Vec::with_capacity(records.len().saturating_mul(excp_stride(flags))); + for term in records { + if pack_excp_point(&mut buf, term, flags).is_err() { + return err(env, atoms::invalid_data()); + } + } + ok_binary(env, &buf) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn gspl_unpack<'a>( + env: Env<'a>, + data: Binary, + sh_rest: u16, + offset: u64, + max_count: u64, +) -> Term<'a> { + unpack_gspl_slice( + env, + data.as_slice(), + sh_rest, + offset as usize, + max_count as usize, + ) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn gspl_pack<'a>(env: Env<'a>, records: Vec, sh_rest: u16) -> Term<'a> { + let mut buf = Vec::with_capacity(records.len().saturating_mul(gspl_stride(sh_rest))); + for term in records { + if pack_gspl_point(&mut buf, term, sh_rest).is_err() { + return err(env, atoms::invalid_data()); + } + } + ok_binary(env, &buf) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn ply_binary_unpack<'a>( + env: Env<'a>, + data: Binary, + types: Vec, + little_endian: bool, + offset: u64, + max_count: u64, +) -> Term<'a> { + unpack_ply_slice( + env, + data.as_slice(), + &types, + little_endian, + offset as usize, + max_count as usize, + ) +} + +#[rustler::nif] +fn spatial_mmap_open<'a>(env: Env<'a>, path: String) -> Term<'a> { + match File::open(&path).and_then(|f| unsafe { Mmap::map(&f) }) { + Ok(mmap) => { + let resource = ResourceArc::new(MappedSpatial { mmap }); + (atoms::ok(), resource).encode(env) + } + Err(_) => err(env, atoms::invalid_data()), + } +} + +#[rustler::nif] +fn spatial_mmap_len(resource: ResourceArc) -> NifResult { + Ok(resource.mmap.len() as u64) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn excp_unpack_mmap<'a>( + env: Env<'a>, + resource: ResourceArc, + flags: u16, + offset: u64, + max_count: u64, +) -> Term<'a> { + unpack_excp_slice( + env, + &resource.mmap, + flags, + offset as usize, + max_count as usize, + ) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn gspl_unpack_mmap<'a>( + env: Env<'a>, + resource: ResourceArc, + sh_rest: u16, + offset: u64, + max_count: u64, +) -> Term<'a> { + unpack_gspl_slice( + env, + &resource.mmap, + sh_rest, + offset as usize, + max_count as usize, + ) +} + +#[rustler::nif(schedule = "DirtyCpu")] +fn ply_binary_unpack_mmap<'a>( + env: Env<'a>, + resource: ResourceArc, + types: Vec, + little_endian: bool, + offset: u64, + max_count: u64, +) -> Term<'a> { + unpack_ply_slice( + env, + &resource.mmap, + &types, + little_endian, + offset as usize, + max_count as usize, + ) +} + +/// Append packed body bytes to a file (chunked encode helper). +#[rustler::nif(schedule = "DirtyCpu")] +fn spatial_append_file<'a>(env: Env<'a>, path: String, data: Binary) -> Term<'a> { + match File::options().append(true).create(true).open(&path) { + Ok(mut file) => match file.write_all(data.as_slice()) { + Ok(()) => atoms::ok().encode(env), + Err(_) => err(env, atoms::invalid_data()), + }, + Err(_) => err(env, atoms::invalid_data()), + } +} diff --git a/test/ex_codecs/spatial/accel_coverage_test.exs b/test/ex_codecs/spatial/accel_coverage_test.exs new file mode 100644 index 0000000..73e5a69 --- /dev/null +++ b/test/ex_codecs/spatial/accel_coverage_test.exs @@ -0,0 +1,551 @@ +defmodule ExCodecs.Spatial.AccelCoverageTest do + use ExUnit.Case, async: true + + alias ExCodecs.Spatial.Accel + alias ExCodecs.Spatial.Codec.{Binary, Gsplat, PLY} + alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} + + defp tmp(ext) do + Path.join( + System.tmp_dir!(), + "ex_codecs_accel_cov_#{System.unique_integer([:positive])}#{ext}" + ) + end + + setup do + if Accel.available?() do + :ok + else + {:skip, "spatial Accel NIF not loaded"} + end + end + + describe "Accel facade" do + test "chunk_size, ply_type_tag, pack/unpack, mmap, and append_file" do + assert Accel.chunk_size() == 4096 + + for t <- [:char, :uchar, :short, :ushort, :int, :uint, :float, :double] do + assert is_integer(Accel.ply_type_tag(t)) + end + + points = [ + Point.new(1.0, 2.0, 3.0, color: {1, 2, 3}, normal: {0.0, 1.0, 0.0}), + Point.new(4.0, 5.0, 6.0) + ] + + assert {:ok, body} = Accel.excp_pack(points, 0b101) + assert {:ok, {decoded, _}} = Accel.excp_unpack(body, 0b101, 0, 10) + assert length(decoded) == 2 + + gs = [ + Gaussian.new({0.0, 0.0, 0.0}, sh: [[0.1, 0.2, 0.3], [0.4, 0.5, 0.6]]), + struct(Gaussian, + position: {1.0, 1.0, 1.0}, + color: {0.2, 0.3, 0.4}, + sh: [0.2, 0.3, 0.4, 0.5, 0.6, 0.7] + ) + ] + + assert {:ok, gbody} = Accel.gspl_pack(gs, 3) + assert {:ok, {gdecoded, _}} = Accel.gspl_unpack(gbody, 3, 0, 10) + assert length(gdecoded) == 2 + + # PLY binary body: 3 floats + ply_body = <<1.0::little-float-32, 2.0::little-float-32, 3.0::little-float-32>> + + assert {:ok, {[[x, y, z]], _}} = + Accel.ply_binary_unpack(ply_body, [:float, :float, :float], :binary_le, 0, 1) + + assert_in_delta x, 1.0, 1.0e-5 + assert_in_delta y, 2.0, 1.0e-5 + assert_in_delta z, 3.0, 1.0e-5 + + assert {:ok, {[[_, _, _]], _}} = + Accel.ply_binary_unpack(ply_body, [:float, :float, :float], true, 0, 1) + + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + {:ok, bin} = Binary.encode(PointCloud.new(points), accel: false) + File.write!(path, bin) + + assert {:ok, ref} = Accel.mmap_open(path) + assert {:ok, len} = Accel.mmap_len(ref) + assert len == byte_size(bin) + + assert {:ok, {mmap_pts, _}} = Accel.excp_unpack_mmap(ref, 0b101, 16, 10) + assert length(mmap_pts) == 2 + + gpath = tmp(".gspl") + on_exit(fn -> File.rm(gpath) end) + {:ok, gbin} = Gsplat.encode(GaussianCloud.new(gs), accel: false) + File.write!(gpath, gbin) + assert {:ok, gref} = Accel.mmap_open(gpath) + assert {:ok, _} = Accel.gspl_unpack_mmap(gref, 3, 18, 10) + + ppath = tmp(".ply") + on_exit(fn -> File.rm(ppath) end) + {:ok, pbin} = PLY.encode(PointCloud.new([Point.new(1, 2, 3)]), ply_format: :binary_le) + File.write!(ppath, pbin) + assert {:ok, pref} = Accel.mmap_open(ppath) + # body offset after a typical small header — unpack may return 0 or more + assert {:ok, {_rows, _}} = + Accel.ply_binary_unpack_mmap(pref, [:float, :float, :float], :binary_le, 0, 1) + + append_path = tmp(".bin") + on_exit(fn -> File.rm(append_path) end) + assert :ok = Accel.append_file(append_path, <<"hello">>) + assert File.read!(append_path) == "hello" + + assert {:error, _} = Accel.mmap_open("/no/such/ex_codecs_mmap") + assert {:error, _} = Accel.append_file("/no/such/dir/x.bin", <<"x">>) + # Short bodies yield an empty chunk (not a NIF error). + assert {:ok, {[], _}} = Accel.excp_unpack(<<"short">>, 0, 0, 1) + assert {:ok, {[], _}} = Accel.gspl_unpack(<<"short">>, 0, 0, 1) + + assert {:error, :invalid_data} = Accel.excp_pack([:not_a_point], 0) + assert {:error, :invalid_data} = Accel.gspl_pack([:not_a_gaussian], 0) + + # Default arity / offset forms + assert {:ok, {_, _}} = Accel.excp_unpack(body, 0b101) + assert {:ok, {_, _}} = Accel.gspl_unpack(gbody, 3) + assert {:ok, {_, _}} = Accel.ply_binary_unpack(ply_body, [:float, :float, :float], :little) + assert {:ok, {_, _}} = Accel.excp_unpack_mmap(ref, 0b101) + assert {:ok, {_, _}} = Accel.gspl_unpack_mmap(gref, 3) + + assert {:ok, {_, _}} = + Accel.ply_binary_unpack_mmap(pref, [:float, :float, :float], :binary_be) + + # Bad resource / args exercise safe/1 rescue paths + assert {:error, _} = Accel.mmap_len(:not_a_resource) + assert {:error, _} = Accel.excp_unpack_mmap(:not_a_resource, 0, 0, 1) + + # Valid struct shape but invalid color for Rust pack → nif_binary error path + bad = struct(Point, x: 1.0, y: 2.0, z: 3.0, color: :nope, normal: nil) + assert {:error, _} = Accel.excp_pack([bad], 0b001) + + bad_g = struct(Gaussian, position: {0.0, 0.0, 0.0}, color: :nope, sh: nil) + assert {:error, _} = Accel.gspl_pack([bad_g], 0) + end + + test "row helpers cover nil SH and nested SH" do + p = Point.new(1, 2, 3, color: {1, 2, 3, 4}, normal: nil) + assert {1.0, 2.0, 3.0, {1, 2, 3, 4}, nil} = Accel.point_to_row(p) + assert %Point{} = Accel.row_to_point(Accel.point_to_row(p)) + + g0 = Gaussian.new({0, 0, 0}, sh: nil) + row0 = Accel.gaussian_to_row(g0, 0) + assert %Gaussian{sh: nil} = Accel.row_to_gaussian(row0, 0) + + g1 = Gaussian.new({0, 0, 0}, color: {0.1, 0.2, 0.3}, sh: [[0.1, 0.2, 0.3], [1.0, 2.0, 3.0]]) + row1 = Accel.gaussian_to_row(g1, 3) + assert %Gaussian{} = Accel.row_to_gaussian(row1, 3) + end + end + + describe "Binary accel: false elixir paths" do + test "encode/decode/stream_encode/stream_decode without Accel" do + cloud = + PointCloud.new([ + Point.new(1.0, 2.0, 3.0, color: {1, 2, 3}, normal: {0.0, 1.0, 0.0}), + Point.new(4.0, 5.0, 6.0, color: {4, 5, 6, 7}) + ]) + + assert {:ok, bin} = Binary.encode(cloud, accel: false) + assert {:ok, decoded} = Binary.decode(bin, accel: false) + assert length(decoded.points) == 2 + + assert [%Point{}, %Point{}] = + Binary.stream_decode(bin, source: :binary, accel: false) |> Enum.to_list() + + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + + assert :ok = + Binary.stream_encode_to_file(cloud.points, path, + schema: [:color, :alpha, :normal], + accel: false + ) + + assert [%Point{}, %Point{}] = + Binary.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + # early halt on IO path + assert [%Point{}] = + Binary.stream_decode(path, source: :file, accel: false) + |> Stream.take(1) + |> Enum.to_list() + end + + test "stream_decode binary accel path errors and truncated file IO path" do + bad_ver = <<"EXCP", 99::little-16, 0::little-16, 1::little-64>> + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(bad_ver, source: :binary, accel: true) |> Enum.to_list() + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(<<"nope">>, source: :binary, accel: true) |> Enum.to_list() + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(<<"nope">>, source: :binary, accel: false) |> Enum.to_list() + + cloud = PointCloud.new([Point.new(1, 2, 3), Point.new(4, 5, 6)]) + {:ok, bin} = Binary.encode(cloud, accel: false) + truncated = binary_part(bin, 0, 20) + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + File.write!(path, truncated) + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + # eof after one full record on IO path + one = <<1.0::little-float-32, 2.0::little-float-32, 3.0::little-float-32>> + + File.write!( + path, + <<"EXCP", 1::little-16, 0::little-16, 2::little-64, one::binary>> + ) + + items = Binary.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + assert match?([%Point{}, {:error, %{reason: :invalid_data}}], items) + + File.write!(path, "") + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + File.write!(path, "EXCP") + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + File.write!( + path, + <<"EXCP", 99::little-16, 0::little-16, 0::little-64>> + ) + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + # bad magic via mmap header path + File.write!(path, <<"XXXX", 1::little-16, 0::little-16, 0::little-64>>) + + assert [{:error, %{reason: :invalid_data}}] = + Binary.stream_decode(path, source: :file, accel: true) |> Enum.to_list() + end + + test "truncated mmap stream and empty cloud" do + assert {:ok, empty} = Binary.encode(PointCloud.new([]), accel: true) + assert {:ok, %{points: []}} = Binary.decode(empty, accel: true) + + assert [] = + Binary.stream_decode(empty, source: :binary, accel: true) |> Enum.to_list() + + # count=2, only one xyz — mmap accel truncated + one = <<1.0::little-float-32, 2.0::little-float-32, 3.0::little-float-32>> + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + + File.write!( + path, + <<"EXCP", 1::little-16, 0::little-16, 2::little-64, one::binary>> + ) + + items = Binary.stream_decode(path, source: :file, accel: true) |> Enum.to_list() + assert Enum.any?(items, &match?({:error, %{reason: :invalid_data}}, &1)) + + # binary stream truncated body + bin = <<"EXCP", 1::little-16, 0::little-16, 2::little-64, one::binary>> + items2 = Binary.stream_decode(bin, source: :binary, accel: true) |> Enum.to_list() + assert Enum.any?(items2, &match?({:error, %{reason: :invalid_data}}, &1)) + end + + test "normalize color branches via elixir encode" do + cloud = + PointCloud.new([ + Point.new(0, 0, 0), + Point.new(1, 1, 1, color: {1, 2, 3}), + Point.new(2, 2, 2, color: {1, 2, 3, 4}) + ]) + + assert {:ok, _} = Binary.encode(cloud, accel: false) + + path = tmp(".excp") + on_exit(fn -> File.rm(path) end) + + # RGB schema with RGBA source exercises normalize_rgb/1 RGBA clause. + assert :ok = + Binary.stream_encode_to_file( + [Point.new(1, 1, 1, color: {9, 8, 7, 6})], + path, + schema: [:color], + accel: false + ) + + assert :ok = + Binary.stream_encode_to_file( + [Point.new(0, 0, 0), Point.new(1, 1, 1, color: {9, 8, 7})], + path, + schema: [:color, :alpha], + accel: false + ) + + assert {:ok, _} = Binary.decode(File.read!(path), accel: false) + + # Missing file on elixir IO open path + assert [{:error, %{reason: :io_error}}] = + Binary.stream_decode("/no/such/ex_codecs_io.excp", source: :file, accel: false) + |> Enum.to_list() + end + end + + describe "Gsplat accel: false elixir paths" do + test "encode/decode/stream without Accel and error edges" do + cloud = + GaussianCloud.new([ + Gaussian.new({1.0, 2.0, 3.0}, opacity: 0.5), + Gaussian.new({0.0, 1.0, 0.0}, + color: {0.2, 0.3, 0.4}, + sh: [[0.2, 0.3, 0.4], [0.1, 0.1, 0.1]] + ) + ]) + + # Empty SH list hits sh_rest_values/1 catch-all list clause (non-empty lists + # match the [_dc | rest] head clause first). + assert {:ok, _} = + Gsplat.encode( + GaussianCloud.new([struct(Gaussian, position: {0.0, 0.0, 0.0}, sh: [])]), + accel: false + ) + + assert {:ok, bin} = Gsplat.encode(cloud, accel: false) + assert {:ok, decoded} = Gsplat.decode(bin, accel: false) + assert length(decoded.gaussians) == 2 + + assert [%Gaussian{}, %Gaussian{}] = + Gsplat.stream_decode(bin, source: :binary, accel: false) |> Enum.to_list() + + path = tmp(".gspl") + on_exit(fn -> File.rm(path) end) + + assert :ok = + Gsplat.stream_encode_to_file(cloud.gaussians, path, + schema: [sh_rest: 3], + accel: false + ) + + assert [%Gaussian{}, %Gaussian{}] = + Gsplat.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + assert [%Gaussian{}] = + Gsplat.stream_decode(path, source: :file, accel: false) + |> Stream.take(1) + |> Enum.to_list() + + assert [{:error, _}] = + Gsplat.stream_decode( + <<"GSPL", 9::little-16, 0::little-16, 1::little-64, 0::little-16>>, + source: :binary, + accel: true + ) + |> Enum.to_list() + + assert [{:error, _}] = + Gsplat.stream_decode(<<"nope">>, source: :binary, accel: true) |> Enum.to_list() + + assert [{:error, _}] = + Gsplat.stream_decode(<<"nope">>, source: :binary, accel: false) |> Enum.to_list() + + File.write!(path, "") + + assert [{:error, _}] = + Gsplat.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + File.write!(path, "GSPL") + + assert [{:error, _}] = + Gsplat.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + File.write!( + path, + <<"GSPL", 9::little-16, 0::little-16, 0::little-64, 0::little-16>> + ) + + assert [{:error, _}] = + Gsplat.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + zeros = :binary.copy(<<0>>, 56) + + File.write!( + path, + <<"GSPL", 1::little-16, 0::little-16, 2::little-64, 0::little-16, zeros::binary>> + ) + + items = Gsplat.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + assert match?([%Gaussian{}, {:error, _}], items) + + items2 = Gsplat.stream_decode(path, source: :file, accel: true) |> Enum.to_list() + assert Enum.any?(items2, &match?({:error, _}, &1)) + + assert {:ok, empty} = Gsplat.encode(GaussianCloud.new([]), accel: true) + assert [] = Gsplat.stream_decode(empty, source: :binary, accel: true) |> Enum.to_list() + + assert [{:error, %{reason: :io_error}}] = + Gsplat.stream_decode("/no/such/ex_codecs_io.gspl", source: :file, accel: false) + |> Enum.to_list() + + # XXXX magic via mmap header + File.write!(path, <<"XXXX", 1::little-16, 0::little-16, 0::little-64, 0::little-16>>) + + assert [{:error, _}] = + Gsplat.stream_decode(path, source: :file, accel: true) |> Enum.to_list() + + # Truncated binary stream (accel) — empty chunk path + short = <<"GSPL", 1::little-16, 0::little-16, 1::little-64, 0::little-16, 0, 1, 2>> + + assert [{:error, %{reason: :invalid_data}}] = + Gsplat.stream_decode(short, source: :binary, accel: true) |> Enum.to_list() + + # Short non-EOF body on IO path (count=1, fewer than stride bytes) + File.write!( + path, + <<"GSPL", 1::little-16, 0::little-16, 1::little-64, 0::little-16, 0, 1, 2, 3, 4>> + ) + + assert [{:error, %{reason: :invalid_data}}] = + Gsplat.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + end + end + + describe "PLY accel: false binary paths" do + test "binary decode and file stream without Accel cover unpack types" do + # Craft a binary PLY with mixed property types so elixir unpack clauses run. + header = """ + ply + format binary_little_endian 1.0 + element vertex 1 + property char c + property uchar uc + property short s + property ushort us + property int i + property uint ui + property float x + property float y + property float z + property double d + end_header + """ + + body = + << + -1::signed-integer-8, + 255::unsigned-integer-8, + -2::little-signed-integer-16, + 3::little-unsigned-integer-16, + -4::little-signed-integer-32, + 5::little-unsigned-integer-32, + 1.25::little-float-32, + 2.5::little-float-32, + 3.75::little-float-32, + 9.0::little-float-64 + >> + + path = tmp(".ply") + on_exit(fn -> File.rm(path) end) + File.write!(path, header <> body) + + assert {:ok, cloud} = PLY.decode(File.read!(path), accel: false) + assert length(cloud.points) == 1 + + assert [%Point{}] = + PLY.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + # big-endian mixed types + header_be = String.replace(header, "little_endian", "big_endian") + + body_be = + << + -1::signed-integer-8, + 255::unsigned-integer-8, + -2::big-signed-integer-16, + 3::big-unsigned-integer-16, + -4::big-signed-integer-32, + 5::big-unsigned-integer-32, + 1.25::big-float-32, + 2.5::big-float-32, + 3.75::big-float-32, + 9.0::big-float-64 + >> + + File.write!(path, header_be <> body_be) + assert {:ok, _} = PLY.decode(File.read!(path), accel: false) + + assert [%Point{}] = + PLY.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + # normal binary_le encode/stream with accel false + cloud2 = PointCloud.new([Point.new(1, 2, 3), Point.new(4, 5, 6)]) + {:ok, bin} = PLY.encode(cloud2, ply_format: :binary_le) + File.write!(path, bin) + + assert [%Point{}, %Point{}] = + PLY.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + assert [%Point{}] = + PLY.stream_decode(path, source: :file, accel: false) + |> Stream.take(1) + |> Enum.to_list() + + # truncated binary body on IO path + File.write!(path, binary_part(bin, 0, byte_size(bin) - 4)) + + items = PLY.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + assert Enum.any?(items, &match?({:error, %{reason: :invalid_data}}, &1)) + + # truncated via mmap accel path + File.write!(path, binary_part(bin, 0, byte_size(bin) - 4)) + items2 = PLY.stream_decode(path, source: :file, accel: true) |> Enum.to_list() + assert Enum.any?(items2, &match?({:error, %{reason: :invalid_data}}, &1)) + + # CR-only newline after end_header (stream IO path uses strip count helper) + File.write!( + path, + "ply\rformat binary_little_endian 1.0\relement vertex 1\rproperty float x\rproperty float y\rproperty float z\rend_header\r" <> + <<1.0::little-float-32, 2.0::little-float-32, 3.0::little-float-32>> + ) + + assert [%Point{}] = + PLY.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + # No newline after end_header — strip_leading_newlines_count/1 catch-all + body = <<1.0::little-float-32, 2.0::little-float-32, 3.0::little-float-32>> + + File.write!( + path, + "ply\nformat binary_little_endian 1.0\nelement vertex 1\nproperty float x\nproperty float y\nproperty float z\nend_header" <> + body + ) + + assert [%Point{}] = + PLY.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + + # First 4096-byte header scan keeps only 1 body byte in leftover so + # take_exact_bytes/3 must read the remainder of the vertex from the file. + prefix = + "ply\nformat binary_little_endian 1.0\nelement vertex 1\nproperty float x\nproperty float y\nproperty float z\n" + + end_h = "end_header\n" + # hdr + 1 body byte == 4096 ⇒ hdr size 4095 + fill_size = 4095 - byte_size(prefix) - byte_size(end_h) + fill = "comment " <> String.duplicate("z", fill_size - 9) <> "\n" + hdr = prefix <> fill <> end_h + assert byte_size(hdr) == 4095 + + File.write!(path, hdr <> body) + + assert [%Point{}] = + PLY.stream_decode(path, source: :file, accel: false) |> Enum.to_list() + end + end +end diff --git a/test/ex_codecs/spatial/accel_property_test.exs b/test/ex_codecs/spatial/accel_property_test.exs new file mode 100644 index 0000000..827fe27 --- /dev/null +++ b/test/ex_codecs/spatial/accel_property_test.exs @@ -0,0 +1,213 @@ +defmodule ExCodecs.Spatial.AccelPropertyTest do + use ExUnit.Case, async: true + use ExUnitProperties + + alias ExCodecs.Spatial.Accel + alias ExCodecs.Spatial.Codec.{Binary, Gsplat, PLY} + alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} + + @moduletag :accel + + setup do + if Accel.available?() do + :ok + else + {:skip, "spatial Accel NIF not loaded"} + end + end + + defp float32 do + # Keep values in a range that survives f32 round-trip without NaN/Inf. + StreamData.float(min: -1.0e5, max: 1.0e5) + |> StreamData.map(fn f -> + <> = <> + x + end) + end + + defp u8, do: StreamData.integer(0..255) + + defp point_generator do + gen all( + x <- float32(), + y <- float32(), + z <- float32(), + mode <- StreamData.member_of([:xyz, :rgb, :rgba, :normal, :rgb_normal]), + r <- u8(), + g <- u8(), + b <- u8(), + a <- u8(), + nx <- float32(), + ny <- float32(), + nz <- float32() + ) do + opts = + case mode do + :xyz -> [] + :rgb -> [color: {r, g, b}] + :rgba -> [color: {r, g, b, a}] + :normal -> [normal: {nx, ny, nz}] + :rgb_normal -> [color: {r, g, b}, normal: {nx, ny, nz}] + end + + Point.new(x, y, z, opts) + end + end + + defp gaussian_generator do + gen all( + x <- float32(), + y <- float32(), + z <- float32(), + r <- float32(), + g <- float32(), + b <- float32(), + opacity <- float32(), + sx <- float32(), + sy <- float32(), + sz <- float32(), + rw <- float32(), + rx <- float32(), + ry <- float32(), + rz <- float32(), + sh_n <- StreamData.integer(0..6), + sh_vals <- StreamData.list_of(float32(), length: sh_n) + ) do + sh = + if sh_n == 0 do + nil + else + [[r, g, b] | Enum.chunk_every(sh_vals, 3)] + end + + Gaussian.new({x, y, z}, + color: {r, g, b}, + opacity: opacity, + scale: {sx, sy, sz}, + rotation: {rw, rx, ry, rz}, + sh: sh + ) + end + end + + defp assert_points_close(a, b) do + assert length(a) == length(b) + + Enum.zip(a, b) + |> Enum.each(fn {p1, p2} -> + assert_in_delta p1.x, p2.x, 1.0e-5 + assert_in_delta p1.y, p2.y, 1.0e-5 + assert_in_delta p1.z, p2.z, 1.0e-5 + assert p1.color == p2.color + assert_normals_close(p1.normal, p2.normal) + end) + end + + defp assert_normals_close(nil, nil), do: :ok + + defp assert_normals_close({a, b, c}, {d, e, f}) do + assert_in_delta a, d, 1.0e-5 + assert_in_delta b, e, 1.0e-5 + assert_in_delta c, f, 1.0e-5 + end + + defp assert_gaussians_close(a, b) do + assert length(a) == length(b) + + Enum.zip(a, b) + |> Enum.each(fn {g1, g2} -> + assert_tuple_close(g1.position, g2.position) + assert_tuple_close(g1.color, g2.color) + assert_in_delta g1.opacity, g2.opacity, 1.0e-5 + assert_tuple_close(g1.scale, g2.scale) + assert_tuple_close(g1.rotation, g2.rotation) + assert_sh_close(g1.sh, g2.sh) + end) + end + + defp assert_tuple_close(t1, t2) do + assert tuple_size(t1) == tuple_size(t2) + + Enum.zip(Tuple.to_list(t1), Tuple.to_list(t2)) + |> Enum.each(fn {a, b} -> assert_in_delta a, b, 1.0e-5 end) + end + + defp assert_sh_close(nil, nil), do: :ok + + defp assert_sh_close(a, b) when is_list(a) and is_list(b) do + assert_tuple_close( + List.to_tuple(List.flatten(a)), + List.to_tuple(List.flatten(b)) + ) + end + + property "EXCP rust decode matches elixir decode" do + check all(points <- StreamData.list_of(point_generator(), min_length: 0, max_length: 40)) do + cloud = PointCloud.new(points) + assert {:ok, bin} = Binary.encode(cloud, accel: false) + assert {:ok, elixir} = Binary.decode(bin, accel: false) + assert {:ok, rust} = Binary.decode(bin, accel: true) + assert_points_close(elixir.points, rust.points) + end + end + + property "EXCP rust pack matches elixir encode body" do + check all(points <- StreamData.list_of(point_generator(), min_length: 1, max_length: 40)) do + assert {:ok, elixir_bin} = Binary.encode(PointCloud.new(points), accel: false) + assert {:ok, rust_bin} = Binary.encode(PointCloud.new(points), accel: true) + assert elixir_bin == rust_bin + end + end + + property "GSPL rust decode matches elixir decode" do + check all(gs <- StreamData.list_of(gaussian_generator(), min_length: 0, max_length: 20)) do + cloud = GaussianCloud.new(gs) + assert {:ok, bin} = Gsplat.encode(cloud, accel: false) + assert {:ok, elixir} = Gsplat.decode(bin, accel: false) + assert {:ok, rust} = Gsplat.decode(bin, accel: true) + assert_gaussians_close(elixir.gaussians, rust.gaussians) + end + end + + property "GSPL rust pack matches elixir encode body" do + check all(gs <- StreamData.list_of(gaussian_generator(), min_length: 1, max_length: 20)) do + assert {:ok, elixir_bin} = Gsplat.encode(GaussianCloud.new(gs), accel: false) + assert {:ok, rust_bin} = Gsplat.encode(GaussianCloud.new(gs), accel: true) + assert elixir_bin == rust_bin + end + end + + property "binary PLY rust decode matches elixir decode" do + check all(points <- StreamData.list_of(point_generator(), min_length: 1, max_length: 30)) do + cloud = PointCloud.new(points) + assert {:ok, bin} = PLY.encode(cloud, ply_format: :binary_le) + assert {:ok, elixir} = PLY.decode(bin, accel: false) + assert {:ok, rust} = PLY.decode(bin, accel: true) + assert_points_close(elixir.points, rust.points) + end + end + + property "EXCP mmap stream_decode matches elixir decode" do + check all(points <- StreamData.list_of(point_generator(), min_length: 1, max_length: 25)) do + assert {:ok, bin} = Binary.encode(PointCloud.new(points), accel: false) + + path = + Path.join( + System.tmp_dir!(), + "ex_codecs_accel_prop_#{System.unique_integer([:positive])}.excp" + ) + + try do + File.write!(path, bin) + assert {:ok, elixir} = Binary.decode(bin, accel: false) + + streamed = + Binary.stream_decode(path, source: :file, accel: true) |> Enum.to_list() + + assert_points_close(elixir.points, streamed) + after + File.rm(path) + end + end + end +end diff --git a/test/ex_codecs/spatial/stream_coverage_test.exs b/test/ex_codecs/spatial/stream_coverage_test.exs index af2fc9b..dd74f53 100644 --- a/test/ex_codecs/spatial/stream_coverage_test.exs +++ b/test/ex_codecs/spatial/stream_coverage_test.exs @@ -2,8 +2,8 @@ defmodule ExCodecs.Spatial.StreamCoverageTest do use ExUnit.Case, async: true alias ExCodecs.Spatial - alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} alias ExCodecs.Spatial.Codec.{Binary, Gsplat, PLY} + alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} alias ExCodecs.Spatial.Stream, as: SpatialStream defp tmp(ext) do From 8ea70760094e69b98ace10844d74ce15da7376d3 Mon Sep 17 00:00:00 2001 From: thanos Date: Sat, 18 Jul 2026 11:37:58 -0400 Subject: [PATCH 4/6] Address v0.2.3 review blockers before publish. Fix NIF availability registration order, drop stale docs from the Hex package, correct streaming/work_factor/from_nif docs, repair livebook demos, and make Accel tests skip cleanly when the spatial NIF is absent. Co-authored-by: Cursor --- CHANGELOG.md | 20 +- docs/architecture.md | 28 +- docs/codec-review.md | 720 --------------- docs/native_architecture.md | 834 ------------------ docs/spatial_formats.md | 10 +- guides/understanding_bzip2.md | 20 +- guides/understanding_spatial_codecs.md | 45 +- lib/ex_codecs.ex | 4 +- lib/ex_codecs/application.ex | 24 +- lib/ex_codecs/spatial.ex | 21 +- livebooks/01_introduction.livemd | 47 +- livebooks/02_compression_fundamentals.livemd | 66 +- livebooks/03_codec_comparison.livemd | 51 +- livebooks/04_building_storage_systems.livemd | 89 +- livebooks/06_spatial_codecs.livemd | 23 +- mix.exs | 2 +- test/ex_codecs/application_test.exs | 46 + .../ex_codecs/spatial/accel_coverage_test.exs | 11 +- .../ex_codecs/spatial/accel_property_test.exs | 8 +- 19 files changed, 331 insertions(+), 1738 deletions(-) delete mode 100644 docs/codec-review.md delete mode 100644 docs/native_architecture.md create mode 100644 test/ex_codecs/application_test.exs diff --git a/CHANGELOG.md b/CHANGELOG.md index 1bcc839..e19f046 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -37,6 +37,22 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - Precompiled NIF checksums must be regenerated when publishing GitHub release artifacts for `0.2.3`. +## [0.2.2] - 2026-07-17 + +### Notes + +- Version number reserved and **superseded; not published** to Hex.pm. The + work intended for 0.2.2 was rolled into 0.2.3 (spatial streaming and Rust + acceleration) instead. Recorded per [Keep a Changelog](https://keepachangelog.com) + so the version-skip is not silent. + +## [0.2.1] - 2026-07-17 + +### Fixed + +- README links and badge URLs corrected following the 0.2.0 spatial release. +- `mix.exs` version bump to 0.2.1. + ## [0.2.0] - 2026-07-16 ### Added @@ -128,7 +144,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - README badges updated with CI, Hex.pm, docs, license, Elixir version, and Coveralls coverage links. -## [0.1.0] - 2025-06-09 +## [0.1.0] - 2026-06-13 ### Added @@ -139,7 +155,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - Blosc2 shuffle support (`:none`, `:byte`). - Rust NIF implementation via `rustler_precompiled`. - Precompiled binaries for macOS (ARM64, x86_64), Linux (glibc, musl, ARM64), Windows (x86_64). -- `ExCodecs.Compression` convenience module (`compress/2`, `decompress/2`). +- `ExCodecs.Compression` convenience module (`compress/3`, `decompress/3`). - Structured error handling with `%ExCodecs.Error{}`. - 154 tests (unit + property-based with StreamData). - 90%+ test coverage. diff --git a/docs/architecture.md b/docs/architecture.md index 51825c5..e5444e2 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -577,25 +577,25 @@ ExCodecs.Error.error(:invalid_data, message: "Data must be a binary") ### NIF error mapping The Rust side returns errors as atoms: `{:error, :compression_failed}`, -`{:error, :invalid_data}`, etc. The Elixir codec modules or the framework -map these into structured errors via `ExCodecs.Error.from_nif/2`: +`{:error, :invalid_data}`, `{:error, :output_limit_exceeded}`, etc. Elixir +codec modules call `ExCodecs.NIF.wrap/2` (or `safe_call/2`) to turn those into +`{:error, %ExCodecs.Error{}}`: ```elixir -def from_nif({:error, reason}, codec) when is_atom(codec) do - {:error, %__MODULE__{ - reason: nif_error_to_atom(reason), - message: "NIF error in codec #{codec}: #{inspect(reason)}", - codec: codec, - details: reason - }} -end +# Typical codec decode path +case ExCodecs.NIF.max_output_size(opts) do + {:ok, max} -> + ExCodecs.NIF.wrap(:zstd, ExCodecs.Native.zstd_decompress(data, max)) -defp nif_error_to_atom(reason) when is_atom(reason), do: reason -defp nif_error_to_atom(_), do: :compression_failed + {:error, _} = err -> + err +end ``` -This creates a boundary: the Rust layer communicates errors as atoms, and the -Elixir layer enriches those atoms into structured errors with context. +`NIF.wrap/2` maps known atoms to structured errors with a default message and +`codec:` field. Unknown atoms still become `%ExCodecs.Error{}` with the raw +atom preserved. This keeps a clear boundary: Rust communicates status as +atoms; Elixir enriches them for callers. ### Error flow diff --git a/docs/codec-review.md b/docs/codec-review.md deleted file mode 100644 index 42bad6c..0000000 --- a/docs/codec-review.md +++ /dev/null @@ -1,720 +0,0 @@ -# ExCodecs Codec Review - -A detailed technical review of each compression codec available in ExCodecs, -covering history, design goals, performance characteristics, ecosystem -adoption, licensing, Rust ecosystem maturity, and tradeoffs. - -## Table of Contents - -- [Zstandard (Zstd)](#zstandard-zstd) -- [LZ4](#lz4) -- [Snappy](#snappy) -- [Bzip2](#bzip2) -- [Blosc2](#blosc2) -- [Comparison Table](#comparison-table) -- [Decision Guide](#decision-guide) - ---- - -## Zstandard (Zstd) - -### History - -Zstandard was created by Yann Collet at Facebook (now Meta) and first released -in 2016. Collet is also the creator of LZ4. Zstd was developed to provide a -compression algorithm that offered both high ratios and high speeds, filling -the gap between LZ4 (fast but lower ratio) and zlib/bzip2 (high ratio but -slow). Facebook deployed Zstd at scale across their infrastructure, and it has -since become one of the most widely adopted modern compression algorithms. - -### Design Goals - -Zstd was designed around four principles: - -1. **High ratio at reasonable speed**: Zstd should compress better than zlib - at comparable or faster speeds, and decompress significantly faster than - zlib across all levels. - -2. **Configurable tradeoff**: The 22 compression levels allow users to choose - any point on the speed/ratio curve, from near-LZ4 speeds to near-bzip2 - ratios. - -3. **Fast decompression**: Regardless of compression level, decompression - speed remains consistently high. This is because decompression is a - simpler, more predictable operation that benefits from optimized rouiting. - -4. **Dictionary compression**: Zstd supports training dictionaries for - compressing many small, structurally similar payloads (e.g., JSON API - responses) where traditional compression performs poorly. - -### Compression Ratio - -Zstd achieves strong compression ratios across its level range: - -| Level | Ratio (Silesia corpus) | Compression Speed | Decompression Speed | -|-------|------------------------|--------------------|---------------------| -| 1 | ~2.8x | ~520 MB/s | ~1300 MB/s | -| 3 | ~3.0x (default) | ~430 MB/s | ~1300 MB/s | -| 9 | ~3.4x | ~140 MB/s | ~1300 MB/s | -| 19 | ~3.7x | ~14 MB/s | ~1300 MB/s | -| 22 | ~3.8x | ~4 MB/s | ~1300 MB/s | - -The decompression speed is nearly constant regardless of compression level. -At level 22, the compressed output is smaller but decompresses just as fast -as level 1 output. - -### Speed Characteristics - -- **Compression**: Highly variable by level. Level 1 approaches LZ4 speeds. - Level 19+ enters the territory of slow but thorough compression suitable - for archival. -- **Decompression**: Extremely fast and consistent. Typical decompression - speeds exceed 1 GB/s, making Zstd an excellent choice for read-heavy - workloads. -- **Memory**: Zstd uses a sliding window for compression (controlled by - `window_log`). Larger windows improve ratio but require more memory. - -### Ecosystem Adoption - -Zstd has seen massive adoption since its release: - -- **Linux kernel**: Zstd is used for kernel and initramfs compression since - Linux 4.14 (2017). -- **File systems**: Btrfs, SquashFS, and F2FS support Zstd compression. -- **Web**: Chrome supports Zstd content encoding (Accept-Encoding: - zstd). CloudFlare and Facebook serve Zstd-compressed responses. -- **Archiving**: tar, rsync, and various backup tools support Zstd. -- **Data formats**: Zstd is a registered compression format in Apache - Parquet, Zarr, and several database systems. -- **GitHub**: Uses Zstd for Git object compression. -- **Docker**: Container image layers can be compressed with Zstd. - -### Licensing - -Zstd is dual-licensed under: - -- **BSD 3-Clause** (patent grant included) -- for general use -- **GPLv2** -- for GPL-compatible projects - -The dual licensing ensures broad compatibility. The BSD license is the -default and covers the vast majority of use cases. Meta has also committed -not to assert patents against Zstd users. - -### Rust Ecosystem Maturity - -The `zstd` crate (used by ExCodecs) is well-maintained and wraps the -reference C implementation via the `zstd-sys` crate. - -| Aspect | Status | -|---------------|--------------------------------------------------| -| Crate | `zstd` v0.13 | -| Backend | C FFI to libzstd (reference implementation) | -| Streaming | Supported (encode/decode all, bulk, streaming) | -| Dictionary | Supported via `zstd::dict` | -| Maintenance | Actively maintained, follows zstd releases | -| Safety | FFI boundary is unsafe; crate provides safe API | - -The use of the C reference implementation via FFI is standard practice in the -Rust ecosystem. Pure Rust Zstd implementations exist but are not mature enough -for production use. The `zstd` crate abstracts the FFI boundary with a safe -API, and Rustler further wraps it in a safe NIF interface. - -### Tradeoffs - -| Pro | Con | -|------------------------------------------------|------------------------------------------------| -| Excellent ratio/speed tradeoff | C dependency via zstd-sys (build complexity) | -| Fast decompression at all levels | Higher memory usage at high levels | -| Wide ecosystem adoption and tooling | NIF binary size larger than pure-Rust codecs | -| 22 compression levels for fine control | Dictionary training requires separate tooling | -| Strong IP position (BSD + patent grant) | Not the fastest compress (LZ4 is faster) | -| Active development and optimization | | - ---- - -## LZ4 - -### History - -LZ4 was created by Yann Collet in 2011 as an evolution of his earlier -LZ4Ultra and FastLZ algorithms. Collet designed LZ4 to prioritize compression -and decompression speed above all else. Its design is rooted in the LZ77 -family of algorithms, with aggressive optimizations for modern CPU -architectures. LZ4 rapidly gained adoption in real-time systems, log -processing, and anywhere that throughput matters more than ratio. - -### Design Goals - -1. **Speed above all**: LZ4 targets compression at over 1 GB/s and - decompression at multi-GB/s speeds. It achieves this through a minimal - match format and branch-free decompression loops. - -2. **Simple format**: The LZ4 block format is deliberately minimal -- a - sequence of token bytes, literal lengths, match lengths, and offsets. - This makes implementations easy to verify and fast to parse. - -3. **Low memory footprint**: LZ4 uses a small hash table for compression - (typically 16 KB) and requires no additional memory for decompression. - -4. **Deterministic performance**: LZ4's speed is consistent and predictable, - with minimal variance across input types. There are no pathological inputs - that cause dramatic slowdowns. - -### Compression Ratio - -LZ4 prioritizes speed over ratio. Typical ratios on the Silesia corpus: - -| Variant | Ratio | Compression Speed | Decompression Speed | -|-----------|--------|--------------------|---------------------| -| LZ4 | ~2.1x | ~700 MB/s | ~3000 MB/s | -| LZ4 HC | ~2.5x | ~40 MB/s | ~3000 MB/s | - -These ratios are lower than Zstd or Bzip2, but the speed advantage is -enormous -- LZ4 decompresses 3-4x faster than Zstd and 10-30x faster than -Bzip2. - -### Speed Characteristics - -- **Compression**: LZ4 block compression is among the fastest general-purpose - algorithms available. LZ4 HC trades compression speed for better ratio - but remains fast to decompress. -- **Decompression**: LZ4 decompression is exceptionally fast (3+ GB/s on - modern hardware) because the format is designed for branch-free, SIMD- - friendly decompression loops. -- **Memory**: Minimal. Compression uses a configurable hash table (default - 16 KB). Decompression requires only the input and output buffers. - -### Ecosystem Adoption - -- **Linux kernel**: LZ4 has been used for kernel and zram compression since - Linux 3.15 (2014). -- **File systems**: Btrfs, SquashFS, and F2FS support LZ4. -- **Databases**: Redis uses LZ4 for list compression. MongoDB supports LZ4 - for document compression. -- **Messaging**: Apache Kafka supports LZ4 compression for topic data. -- **Networking**: OpenVPN and various VPN products use LZ4 for in-flight - compression. -- **Logging**: Various log aggregation tools use LZ4 for real-time - compression of log streams. - -### Licensing - -LZ4 is licensed under **BSD 2-Clause** (simplified BSD license). This is a -permissive license with no patent concerns. - -### Rust Ecosystem Maturity - -ExCodecs uses `lz4_flex`, a pure Rust LZ4 implementation: - -| Aspect | Status | -|---------------|--------------------------------------------------| -| Crate | `lz4_flex` v0.11 | -| Backend | Pure Rust (no C FFI) | -| Format | LZ4 block and frame format | -| Safety | 100% safe Rust | -| Maintenance | Actively maintained | - -The choice of `lz4_flex` over the C-based `lz4` crate was deliberate: - -1. **No C dependency**: Pure Rust avoids build complexity and cross- - compilation issues. This is especially important for `rustler_precompiled` - targets. -2. **Safety**: No unsafe FFI boundary. The entire compression and - decompression path is safe Rust. -3. **Simplicity**: The pure Rust crate has fewer build dependencies and - produces smaller binaries. -4. **Performance**: `lz4_flex` achieves comparable speeds to the C reference - implementation on modern hardware. - -### Tradeoffs - -| Pro | Con | -|------------------------------------------------|------------------------------------------------| -| Extremely fast compression and decompression | Lower compression ratio | -| Pure Rust implementation (lz4_flex) | Not suitable for archival or storage | -| Minimal memory usage | Large inputs can produce marginal ratios | -| Deterministic, predictable speed | LZ4 HC (higher ratio) is much slower | -| Simple, well-understood format | No streaming support in ExCodecs yet | -| No C dependencies for cross-compilation | Frame format not exposed in current API | - ---- - -## Snappy - -### History - -Snappy was originally created by Google under the name "Zippy" and was -renamed and open-sourced as Snappy in 2011. It was designed for internal use -in Google's infrastructure -- compressing data for Bigtable, MapReduce, and -inter-process communication. Snappy's design philosophy centers on maximum -throughput with minimal CPU overhead, targeting the use case where data is -compressed for transient transport and decompressed immediately on the other -end. - -### Design Goals - -1. **Maximum throughput**: Snappy targets compression speeds exceeding - 500 MB/s and decompression speeds exceeding 1.5 GB/s. It sacrifices - compression ratio to achieve this. - -2. **Minimal overhead**: The Snappy format adds very little metadata. - The overhead for incompressible data is minimal (typically 5-6 bytes - per 32 KB block plus the literals). - -3. **Stability**: Snappy's format and behavior are deterministic. The same - input always produces the same output. This makes it suitable for - content-addressable storage where bit-exact reproduction is required. - -4. **Simplicity**: Snappy has no configuration options. There is one - compression strategy, one decompression path, and one output format. - This makes it easy to implement correctly and fast to verify. - -### Compression Ratio - -Snappy achieves the lowest compression ratios among the ExCodecs codecs, -which is the expected tradeoff for its speed: - -| Metric | Typical Value (Silesia) | -|----------------------------|-------------------------| -| Compression ratio | ~2.0x | -| Compression speed | ~500-600 MB/s | -| Decompression speed | ~1500-2000 MB/s | - -For structured data (JSON, protocol buffers), Snappy typically achieves -2.0-2.5x compression. For random binary data, it may produce output larger -than the input (though the format handles this gracefully). - -### Speed Characteristics - -- **Compression**: Extremely fast and consistent. Snappy uses a simple hash - table and short match encoding. There are no expensive computations. -- **Decompression**: Among the fastest decompressors available. The format - is designed for SIMD-friendly, branch-predictable decompression. -- **Memory**: Very small. Compression uses a ~32 KB hash table. - Decompression is single-pass with no additional allocation beyond the - output buffer. - -### Ecosystem Adoption - -- **Google infrastructure**: Bigtable, MapReduce, Protocol Buffers (as an - option), and many internal Google systems. -- **Apache projects**: Hadoop, Cassandra, and various Apache databases - support Snappy compression. -- **Data formats**: The Snappy framing format is used in Parquet, ORC, and - other columnar storage formats. -- **Networking**: Various RPC frameworks support Snappy for compressing - request/response payloads. - -### Licensing - -Snappy is licensed under **BSD 3-Clause**. Google holds the copyright and -has made no patent claims related to Snappy. - -### Rust Ecosystem Maturity - -ExCodecs uses the `snap` crate: - -| Aspect | Status | -|---------------|--------------------------------------------------| -| Crate | `snap` v1.1 | -| Backend | Pure Rust | -| Format | Raw and framing format supported | -| Safety | Mostly safe Rust; some unsafe for SIMD paths | -| Maintenance | Actively maintained | - -The `snap` crate implements both the raw (block) format and the framing -format. ExCodecs uses only the raw format (`snap::raw::Encoder` and -`snap::raw::Decoder`), which is the simpler and faster of the two. - -### Tradeoffs - -| Pro | Con | -|------------------------------------------------|------------------------------------------------| -| Extremely fast compression and decompression | Lowest compression ratio in the set | -| No configuration needed (one mode) | No tunable parameters at all | -| Simple, well-tested format | Not suitable for archival or storage | -| Deterministic output (same input = same output)| May expand random/incompressible data | -| Pure Rust implementation available | Less flexible than LZ4 (no HC variant) | -| Very small memory footprint | Deprecated at Google in favor of Zstd | - -**Note on Snappy's future**: Google has largely moved to Zstd for internal use. -Snappy remains widely deployed and supported, but it is effectively in -maintenance mode. New projects should consider Zstd or LZ4 instead, unless -they need Snappy for compatibility with existing data formats. - ---- - -## Bzip2 - -### History - -Bzip2 was created by Julian Seward in 1996 and released as open source. It -was one of the first widely available compression algorithms to use the -Burrows-Wheeler Transform (BWT), a technique discovered by Michael Burrows -and David Wheeler in 1994. Seward's insight was that BWT followed by -move-to-front coding and Huffman coding could achieve compression ratios -competitive with PPM (Prediction by Partial Matching) algorithms at much -higher speeds. Bzip2 quickly became a standard for software distribution and -archival on Unix systems. - -### Design Goals - -1. **High compression ratio**: Bzip2 targets compression ratios competitive - with the best available algorithms. It consistently achieves among the - highest ratios of any lossless general-purpose compressor. - -2. **Recoverability**: Bzip2's block-based structure allows partial - decompression. If a compressed file is damaged, the blocks before the - damage can still be recovered. - -3. **Simplicity of interface**: Bzip2 has a simple API with one primary - parameter -- block size (1-9). This makes it easy to integrate and - difficult to misconfigure. - -4. **Stability**: The bzip2 format has been stable since 2000. Compressed - data from 2000 can still be decompressed today. - -### Compression Ratio - -Bzip2 achieves the highest compression ratios among ExCodecs' general-purpose -codecs: - -| Block Size | Ratio (Silesia) | Compression Speed | Decompression Speed | -|------------|------------------|--------------------|---------------------| -| 1 | ~2.9x | ~15 MB/s | ~30 MB/s | -| 5 | ~3.2x | ~10 MB/s | ~28 MB/s | -| 9 (default)| ~3.3x | ~8 MB/s | ~25 MB/s | - -These ratios are higher than Zstd at level 9-19 but come at a significant -speed cost. Bzip2 is roughly 50-100x slower at decompression than LZ4. - -### Speed Characteristics - -- **Compression**: Slow. Bzip2's BWT and Huffman coding passes are - computationally expensive. At block size 9, compression speeds are under - 10 MB/s. -- **Decompression**: Also slow compared to modern algorithms. The BWT - inverse transform requires significant computation per block. -- **Memory**: Increases with block size. At block size 9, compression - requires approximately 8 MB of memory for the BWT workspace. This is - significantly more than LZ4 (16 KB) or Zstd (varies by level). -- **Blocking**: Bzip2 processes data in 100 KB - 900 KB blocks (controlled - by block size). This provides natural boundaries for parallel processing - (though ExCodecs does not currently expose parallel decompression). - -### Ecosystem Adoption - -- **Software distribution**: Many Linux distributions distribute source - tarballs as `.tar.bz2`. The kernel was historically distributed as - `.tar.bz2` before moving to Zstd. -- **Archival**: Bzip2 is widely used for long-term storage where ratio - matters more than speed. -- **Unix standard**: `bzip2` has been a standard Unix tool for decades. - Every major Linux distribution includes it. -- **Data exchange**: Some scientific data formats use bzip2 for compressing - large datasets. - -### Licensing - -Bzip2 uses a **BSD-like license** (based on the BSD 4-Clause license with an -additional advertising clause). The license is permissive and similar in -spirit to BSD, though the advertising clause is somewhat unusual. - -### Rust Ecosystem Maturity - -ExCodecs uses the `bzip2` crate: - -| Aspect | Status | -|---------------|--------------------------------------------------| -| Crate | `bzip2` v0.4 | -| Backend | C FFI to libbz2 (reference implementation) | -| Streaming | Supported (BzEncoder, BzDecoder) | -| Safety | Unsafe FFI boundary; safe Rust API wrapper | -| Maintenance | Maintained, follows libbz2 releases | - -The `bzip2` crate wraps the reference C implementation (`libbz2`) via FFI. -There is no mature pure Rust bzip2 implementation; the BWT and Huffman coding -are complex enough that a from-scratch Rust implementation would require -significant engineering and verification effort. - -ExCodecs uses the streaming API (`bzip2::write::BzEncoder` and -`bzip2::read::BzDecoder`) even for one-shot compression/decompression. -This is because the streaming API handles the block-based nature of bzip2 -correctly and allows the same code path to be extended for streaming support -in the future. - -### Tradeoffs - -| Pro | Con | -|------------------------------------------------|------------------------------------------------| -| Highest compression ratio (general purpose) | Very slow compression and decompression | -| Stable format (unchanged since 2000) | High memory usage at large block sizes | -| Block-based (partial recovery on corruption) | C dependency via bzip2-sys | -| Simple API (one parameter: block size) | No streaming support in ExCodecs yet | -| Widely available utility (bzip2 command) | Not suitable for real-time applications | -| Strong data integrity checks | Single-threaded (no parallel compression) | - ---- - -## Blosc2 - -### History - -Blosc2 was created by Francesc Alted starting in 2021, building on the -original Blosc (created in 2010). Blosc was designed as a meta-compressor -for numerical data in scientific computing, particularly for the PyTables -and bcolz projects. The "Blosc" name comes from "Blocking and Shuffling -Optimized Compression." Blosc2 extended the original format with a new -header structure, support for more internal compressors, and improved -multithreading. - -Blosc2's key insight is that numerical arrays (float64, int32, etc.) compress -much better when the bytes are rearranged before compression. By transposing -an array so that similar-valued bytes are adjacent, standard compressors -like LZ4 and Zstd achieve dramatically better ratios on structured data. - -### Design Goals - -1. **Meta-compression**: Blosc2 is not a compression algorithm itself. It is - a framework that applies byte/bit shuffling followed by an internal - compressor (LZ4, Zstd, Snappy, BloscLZ, or zlib). The caller chooses - both the compressor and the shuffle strategy. - -2. **Array-optimized**: Blosc2 is designed for data whose length is a - multiple of `typesize` -- the size of each element in the array. When - `typesize` is set correctly and shuffle is enabled, compression ratios - on numerical arrays can improve by 2-10x. - -3. **Zero-overhead passthrough**: When compression produces output larger - than the input (common with small or random data), Blosc2 stores the - data uncompressed with minimal overhead -- just the 16-byte header. - -4. **Multithreading**: The C-Blosc2 library supports multi-threaded - compression and decompression via a thread pool. (Note: ExCodecs' - pure Rust implementation currently uses single-threaded mode.) - -5. **Self-describing format**: The Blosc2 header includes the compressor - type, compression level, shuffle mode, and typesize, allowing - decompression without external metadata. - -### Compression Ratio - -Blosc2's ratio depends heavily on the data type and shuffle setting: - -| Data Type | Shuffle | Internal | Ratio (typical) | -|------------------------|-----------|----------|-----------------| -| Float64 array | byte | LZ4 | 4-10x | -| Float64 array | byte | Zstd | 6-15x | -| Float64 array | none | LZ4 | 1.5-3x | -| Float64 array | none | Zstd | 2-5x | -| Random binary | none | any | ~1.0x (passthrough) | -| JSON text | none | Zstd | 2-3x | - -The shuffle step is what makes Blosc2 exceptional for typed arrays. Consider -an array of 64-bit floats: `[1.0, 2.0, 3.0, ...]`. In memory, this looks like: - -``` -Byte layout (8 bytes per float, no shuffle): -[3f f0 00 00 00 00 00 00] [40 00 00 00 00 00 00 00] [40 08 00 00 00 00 00 00] - ^-- float 1.0 ----------- ^-- float 2.0 ----------- ^-- float 3.0 ---------- - -Byte layout (after byte shuffle, grouping bytes by position): -[3f 40 40 ...] [f0 00 08 ...] [00 00 00 ...] [00 00 00 ...] [00 00 00 ...] [00 00 00 ...] [00 00 00 ...] [00 00 00 ...] - ^-- byte 0 ^-- byte 1 ^-- byte 2 ... (repeated zeros = excellent compression) -``` - -After shuffle, runs of similar bytes become adjacent, and compressors like -LZ4 and Zstd achieve dramatically better ratios on the run-length-encoded -zeros. - -### Speed Characteristics - -- **Compression**: Fast when using LZ4 or BloscLZ as the internal compressor. - The shuffle step adds some overhead but is typically cheap relative to the - compression itself. -- **Decompression**: Fast. The unshuffle step is cheap (O(n) byte - transposition), and the internal decompressor is fast. -- **Memory**: Blosc2 processes data in blocks (configurable via - `blocksize`). Smaller blocks reduce memory usage but may reduce ratio. -- **Threading**: The C-Blosc2 library supports multi-threaded compression - and decompression. ExCodecs' pure Rust implementation is currently - single-threaded, though the `numthreads` parameter is accepted for forward - compatibility. - -### Ecosystem Adoption - -- **Scientific computing**: Blosc2 is the primary compression format for - PyTables, bcolz, and the newer python-blosc2 library. -- **Data formats**: The HDF5 library can use Blosc2 as a compression filter. - Zarr supports Blosc2 as a compressor. -- **NumPy**: The blosc2 Python package provides NumPy-aware compression. -- **C-Blosc2**: The reference C library is used across the scientific Python - ecosystem. - -### Licensing - -Blosc2 is licensed under **BSD 3-Clause**. - -### Rust Ecosystem - -ExCodecs uses a **pure Rust implementation** of the Blosc2 format rather than -binding to the C-Blosc2 library. This was a deliberate design decision. - -| Aspect | Status | -|---------------|--------------------------------------------------| -| Implementation | Pure Rust in blosc2_codec.rs | -| Internal codecs| LZ4 (lz4_flex), Zstd (zstd), Snappy (snap) | -| Shuffle | Byte shuffle/unshuffle, bit shuffle/unshuffle | -| Threading | Single-threaded (numthreads parameter accepted) | -| Coverage | Core compress/decompress, all major options | - -### Why pure Rust instead of C-Blosc2 FFI? - -1. **No C dependency**: C-Blosc2 depends on multiple internal libraries - (libzstd, liblz4, etc.) and has a complex build system. Binding to it - via FFI would add significant complexity to the NIF build process and - complicate `rustler_precompiled` distribution. - -2. **Code reuse**: The internal compress/decompress functions already exist - in the NIF (LZ4 via `lz4_flex`, Zstd via `zstd`, Snappy via `snap`). - Blosc2's meta-compression pattern just calls these with shuffled data. - Reusing the existing Rust crates avoids code duplication. - -3. **Format simplicity**: The Blosc2 header is only 16 bytes with a - well-documented structure. Parsing and constructing the header in Rust - takes approximately 100 lines, far less effort than binding C-Blosc2. - -4. **Build reproducibility**: A pure Rust implementation compiles - deterministically with the same toolchain. No system libraries, no - pkg-config, no cmake. - -5. **The shuffle is the value**: The primary benefit of Blosc2 is the byte - and bit shuffle. The header format and passthrough behavior are trivial - to implement. The shuffle routines are approximately 30-50 lines of Rust - each. - -**Limitations of the pure Rust approach**: - -- No multi-threaded compression (the `numthreads` parameter is accepted but - ignored). This can be added in a future release using Rayon or a similar - Rust parallelism library. -- The bit shuffle implementation is a simplified version that handles common - cases correctly but may differ from C-Blosc2 for edge cases. -- Some advanced C-Blosc2 features (lazy decompression, filters beyond - shuffle, frames) are not yet supported. - -### Tradeoffs - -| Pro | Con | -|------------------------------------------------|------------------------------------------------| -| Dramatically better ratio on typed arrays | Poor ratio on random/non-structured data | -| Self-describing format (header includes meta) | Pure Rust impl lacks multithreading | -| Supports multiple internal compressors | Not a drop-in replacement for C-Blosc2 data | -| Zero-overhead passthrough on incompressible | Bit shuffle implementation simplified | -| Pure Rust implementation (no C deps) | numthreads parameter currently ignored | -| Reuses existing Rust codec crates | Some C-Blosc2 features not yet supported | -| Excellent for numerical/array data | Requires typesize knowledge for best results | - ---- - -## Comparison Table - -| Property | Zstd | LZ4 | Snappy | Bzip2 | Blosc2 | -|-------------------|-----------------|-----------------|-----------------|-----------------|----------------------| -| **Category** | General-purpose | General-purpose | General-purpose | General-purpose | Meta-compressor | -| **Ratio** | High (2.8-3.8x) | Low (2.1x) | Low (2.0x) | Very high (3.3x)| Variable (1-15x) | -| **Compress Speed**| Fast-slow* | Very fast | Very fast | Slow | Fast (depends on internal) | -| **Decomp Speed** | Very fast | Extremely fast | Extremely fast | Slow | Fast (depends on internal) | -| **Memory (comp)** | 1-128 MB | 16 KB | 32 KB | 1-8 MB | Block-size dependent | -| **Configurability**| 22 levels | 1-16 (HC variant)| None | Block size 1-9 | cname, clevel, shuffle, typesize, blocksize | -| **Streaming?** | Yes | No (block only) | No (raw only) | Yes (block-based)| Yes (block-based) | -| **Rust Backend** | zstd (C FFI) | lz4_flex (pure) | snap (pure) | bzip2 (C FFI) | Pure Rust w/ existing crates | -| **Best For** | Balanced use | Real-time/lowest latency | Short-lived data | Archival/storage | Numerical arrays | -| **License** | BSD/GPL dual | BSD 2-Clause | BSD 3-Clause | BSD-like | BSD 3-Clause | -| **ExCodecs Default**| Level 3 | Level 1 | N/A | Block 9 / WF 30 | LZ4, level 5, byte shuffle | - -*Zstd compression speed varies from very fast (level 1, ~520 MB/s) to slow -(level 22, ~4 MB/s). Decompression is consistently fast (~1.3 GB/s). - -### Speed Comparison (Approximate, Silesia Corpus) - -``` -Compression Speed (MB/s, higher is better) -LZ4 ████████████████████████████████████████ ~700 -Snappy ██████████████████████████████████ ~550 -Zstd-1 ███████████████████████████ ~520 -Zstd-3 █████████████████████████ ~430 -Blosc2 ██████████████████████ ~350* -Bzip2-9 ██ ~8 - -Decompression Speed (MB/s, higher is better) -LZ4 ████████████████████████████████████████ ~3000 -Snappy ████████████████████████████ ~1700 -Zstd ██████████████████████ ~1300 -Blosc2 ██████████████████ ~1100* -Bzip2-9 █ ~25 - -Compression Ratio (higher is better) -Bzip2-9 █████████████████████████████████ ~3.3x -Zstd-22 ████████████████████████████████ ~3.8x -Zstd-3 ██████████████████████████ ~3.0x -Zstd-1 ████████████████████████ ~2.8x -LZ4 ████████████████████ ~2.1x -Snappy ██████████████████ ~2.0x - -* Blosc2 speeds depend heavily on internal compressor and data typesize. - Numbers shown use LZ4 as internal compressor with byte shuffle on float64. -``` - -### When to Use Each Codec - -| Use Case | Recommended Codec | Reason | -|---------------------------------------|--------------------|----------------------------------------| -| General-purpose compression | Zstd level 3 | Best ratio/speed tradeoff | -| Real-time/low-latency compression | LZ4 | Fastest compression and decompression | -| Short-lived data (RPC, caching) | Snappy | Minimal overhead, deterministic output | -| Archival/storage | Bzip2 | Highest ratio (or Zstd level 19-22) | -| Numerical arrays (float64, int32) | Blosc2 | Shuffle dramatically improves ratio | -| Content-addressable storage | Zstd | Deterministic output at same level | -| Maximum decompression speed | LZ4 | 3+ GB/s decompression | -| Minimum memory usage | LZ4 or Snappy | 16-32 KB compression buffer | -| Streaming compression | Zstd | Native streaming API | -| Binary data with known element size | Blosc2 | Shuffle exploit element structure | - ---- - -## Decision Guide - -``` - What are you compressing? - | - +------------+-----------+ - | | - Numerical arrays? General binary data? - (float64, int32, etc.) (text, JSON, arbitrary) - | | - Use Blosc2 What matters more? - with byte shuffle | - +---------+---------+ - | | - Speed? Ratio? - | | - Use LZ4 How much speed - (1 GB/s+) are you willing - to sacrifice? - | - +-------+-------+ - | | - Some A lot - | | - Use Zstd-3 Use Bzip2 - (fast + good (best ratio, - ratio) slow speed) -``` - -For most use cases, **Zstd at level 3** is the recommended default. It -provides a strong compression ratio at high speed, and its decompression -performance is excellent. Switch to LZ4 only when latency is critical and -the ratio penalty is acceptable. Switch to Bzip2 only for archival use -where decompression speed is irrelevant. Use Blosc2 when your data is -numerical arrays and you can specify the element size. \ No newline at end of file diff --git a/docs/native_architecture.md b/docs/native_architecture.md deleted file mode 100644 index f454020..0000000 --- a/docs/native_architecture.md +++ /dev/null @@ -1,834 +0,0 @@ -# ExCodecs Native/NIF Architecture - -The Rust NIF layer is the engine of ExCodecs. Every compression and -decompression operation passes through a Rustler NIF boundary into compiled -Rust code. This document describes how that boundary works, how binaries are -handled, how errors propagate, and how the system stays safe under load. - -## Table of Contents - -- [Architecture Overview](#architecture-overview) -- [Rust NIF Design](#rust-nif-design) -- [Precompiled Distribution](#precompiled-distribution) -- [Error Mapping](#error-mapping) -- [Scheduler Considerations](#scheduler-considerations) -- [Binary Handling](#binary-handling) -- [Safety Considerations](#safety-considerations) -- [Codec Implementation Patterns](#codec-implementation-patterns) -- [NIF Call Flow](#nif-call-flow) - ---- - -## Architecture Overview - -``` - +-------------------+ +--------------------+ +-------------------+ - | Elixir Client | | ExCodecs.Native | | Rust NIF | - | | | (Rustler module) | | (cdylib crate) | - | ExCodecs.encode( | --+->| zstd_compress/2 | --+->| zstd_codec:: | - | :zstd, data, [] | | snappy_compress/1 | | | zstd_compress() | - | ) | | blosc2_compress/7 | | | | - +-------------------+ +--------------------+ +-------------------+ - | | - | :erlang.nif_error/2 | Rust lib crates - | (fallback) | - v v - Process crash +-------------------+ - | zstd crate | - | lz4_flex crate | - | snap crate | - | bzip2 crate | - | (pure Rust blosc2)| - +-------------------+ -``` - -The architecture has three layers: - -1. **Elixir codec modules** -- Validate options, delegate to the Native module. -2. **ExCodecs.Native** -- The Rustler-generated boundary. Defines NIF function - stubs with `:erlang.nif_error/2` fallbacks. Maps directly to Rust functions. -3. **Rust codec modules** -- Each in its own `.rs` file, calling Rust crate - implementations and converting results into BEAM terms. - ---- - -## Rust NIF Design - -### Crate configuration - -The Rust crate is a `cdylib` targeting the BEAM NIF interface: - -```toml -[package] -name = "ex_codecs_native" -version = "0.1.0" -edition = "2021" -rust-version = "1.77" - -[lib] -name = "ex_codecs_native" -path = "src/lib.rs" -crate-type = ["cdylib"] -``` - -The `cdylib` crate type produces a shared library that the BEAM can load as a -NIF. This is the only crate type Rustler supports. - -### Rustler initialization - -The NIF module is registered in `lib.rs` with the `rustler::init!` macro: - -```rust -rustler::init!( - "Elixir.ExCodecs.Native", - [ - zstd_codec::zstd_compress, - zstd_codec::zstd_decompress, - lz4_codec::lz4_compress, - lz4_codec::lz4_decompress, - snappy_codec::snappy_compress, - snappy_codec::snappy_decompress, - bzip2_codec::bzip2_compress, - bzip2_codec::bzip2_decompress, - blosc2_codec::blosc2_compress, - blosc2_codec::blosc2_decompress, - codec_versions, - ], - load = load -); - -fn load(_env: Env, _term: Term) -> bool { - true -} -``` - -The first argument (`"Elixir.ExCodecs.Native"`) must match the Elixir module -name exactly. The second argument is the list of NIF functions exposed to the -BEAM. The `load` callback runs once when the NIF is loaded and returns `true` -to indicate success. - -### Atom definitions - -Atoms are defined centrally in `atoms.rs` and shared across all codec modules: - -```rust -use rustler::atoms; - -atoms! { - ok, - error, - unsupported_codec, - codec_unavailable, - invalid_data, - invalid_options, - compression_failed, - decompression_failed, - nif_not_loaded, -} -``` - -The `atoms!` macro pre-allocates these atoms at NIF load time, avoiding -runtime atom table lookups. Each codec module imports these atoms with -`use crate::atoms` to construct return tuples. - -The atoms mirror the `error_reason` type defined in `ExCodecs.Error`: - -```elixir -@type error_reason :: - :unsupported_codec - | :codec_unavailable - | :invalid_data - | :invalid_options - | :compression_failed - | :decompression_failed - | :nif_not_loaded -``` - -This one-to-one mapping ensures that NIF error atoms are always valid -`error_reason` values. If the Rust side returns an unrecognized atom, the -Elixir error mapper falls back to `:compression_failed`. - -### NIF function stubs on the Elixir side - -The `ExCodecs.Native` module uses Rustler to generate the actual NIF binding: - -```elixir -defmodule ExCodecs.Native do - use Rustler, - otp_app: :ex_codecs, - crate: :ex_codecs_native, - mode: :release -end -``` - -Each function also has a fallback that executes when the NIF is not loaded: - -```elixir -def zstd_compress(_data, _level), do: :erlang.nif_error(:nif_not_loaded) -def zstd_decompress(_data), do: :erlang.nif_error(:nif_not_loaded) -def lz4_compress(_data, _level), do: :erlang.nif_error(:nif_not_loaded) -def lz4_decompress(_data), do: :erlang.nif_error(:nif_not_loaded) -# ... etc -``` - -The `:erlang.nif_error/1` call raises an error at runtime when the NIF is -absent. The application startup detects this via `function_exported?/3`: - -```elixir -defp nif_loaded? do - function_exported?(ExCodecs.Native, :zstd_compress, 2) -rescue - _ -> false -end -``` - -If the NIF is not loaded, all codecs are registered as unavailable rather -than causing runtime crashes. - ---- - -## Precompiled Distribution - -### The rustler_precompiled pipeline - -ExCodecs uses the `rustler_precompiled` package to distribute pre-compiled -NIF binaries. This eliminates the need for a Rust compiler on the user's -machine. - -The pipeline: - -``` - Developer machine CI / Release User machine -+--------------------+ +-----------------------+ +------------------+ -| Write Rust code | | Build NIF for each | | mix deps.get | -| in native/ | | target triple: | | | -| | --push-> | | | rustler_precompiled -| mix rustler_precompiled. | aarch64-apple-darwin | | downloads the | -| build | | x86_64-apple-darwin | | matching .so | -| | | x86_64-linux-gnu | | | -| produces .so files | | x86_64-linux-musl | | No Rust compiler | -| for local target | | aarch64-linux-gnu | | needed | -| | | aarch64-linux-musl | +------------------+ -+--------------------+ | x86_64-windows-msvc | - +-----------------------+ -``` - -### Target configuration - -The target triples are defined in `mix.exs`: - -```elixir -defp rustler_precompiled do - [ - targets: [ - "aarch64-apple-darwin", - "x86_64-apple-darwin", - "x86_64-unknown-linux-gnu", - "x86_64-unknown-linux-musl", - "aarch64-unknown-linux-gnu", - "aarch64-unknown-linux-musl", - "x86_64-pc-windows-msvc" - ], - mode: :release, - nif_versions: ["2.17"] - ] -end -``` - -The `nif_versions: ["2.17"]` setting specifies the NIF API version. NIF -version 2.17 corresponds to OTP 26+ and is binary-compatible with all later -OTP versions that support NIF 2.17. - -### Release optimization - -The `Cargo.toml` release profile is optimized for binary size and speed: - -```toml -[profile.release] -opt-level = 3 # Maximum optimization -lto = true # Link-time optimization across crates -codegen-units = 1 # Single codegen unit for better optimization -strip = true # Strip debug symbols from binary -``` - -LTO (Link-Time Optimization) allows the compiler to optimize across crate -boundaries, which can significantly reduce binary size when multiple codec -implementations are linked together. The single codegen unit gives the -optimizer more context, and stripping removes unnecessary debug info. - ---- - -## Error Mapping - -### The NIF error protocol - -All NIF functions return one of two tuple shapes: - -``` -{:ok, binary()} -- Success -{:error, atom()} -- Failure -``` - -On the Rust side, this is constructed using the pre-allocated atoms: - -```rust -// Success -(atoms::ok(), Binary::new(env, output.as_slice())).encode(env) - -// Failure -(atoms::error(), atoms::compression_failed()).encode(env) -``` - -### Error flow diagram - -``` - Rust NIF Elixir -+--------------------------------------------+ +---------------------------+ -| match zstd::bulk::compress(data, level) { | | | -| Ok(compressed) => | | {:ok, compressed} | -| (atoms::ok(), binary).encode(env) |---> | | -| | | | -| Err(_) => | | {:error, :compression_ | -| (atoms::error(), atoms::compression_ | | failed} | -| _failed()).encode(env) |---> | | -| } | | | -+--------------------------------------------+ | ExCodecs.Error.from_nif() | - | maps atom -> %Error{} | - +---------------------------+ -``` - -### Mapping in detail - -The Elixir `ExCodecs.Error.from_nif/2` function enriches bare NIF error atoms -with context: - -```elixir -def from_nif({:error, reason}, codec) when is_atom(codec) do - {:error, %__MODULE__{ - reason: nif_error_to_atom(reason), - message: "NIF error in codec #{codec}: #{inspect(reason)}", - codec: codec, - details: reason - }} -end - -defp nif_error_to_atom(reason) when is_atom(reason), do: reason -defp nif_error_to_atom(_), do: :compression_failed -``` - -The `nif_error_to_atom` fallback ensures that even if the NIF returns an -unexpected atom or non-atom value, the error maps to a known reason. - -### Error categories - -| NIF atom | Elixir error reason | When it occurs | -|------------------------|-------------------------|---------------------------------------| -| `compression_failed` | `:compression_failed` | Compression algorithm returns error | -| `decompression_failed` | `:decompression_failed` | Decompression algorithm returns error | -| `invalid_data` | `:invalid_data` | Input is malformed (e.g., bad header) | -| `invalid_options` | `:invalid_options` | Options are out of range | - -Note: option validation happens at the Elixir layer before the NIF is called. -The NIF should never receive invalid options. However, `invalid_options` is -defined as an atom for defense-in-depth. - -### Which layer validates what - -``` -+------------------+ +------------------+ +------------------+ -| ExCodecs module | | Codec module | | Rust NIF | -| (public API) | | (per-algorithm) | | (implementation) | -+------------------+ +------------------+ +------------------+ -| Validates: | | Validates: | | Validates: | -| - codec exists | | - level ranges | | - data length | -| - data is binary| | - option types | | - internal | -| - opts is list | | - option values | | constraints | -+------------------+ +------------------+ +------------------+ -``` - -The Elixir layers handle structural validation (types, ranges). The Rust layer -handles algorithmic validation (malformed input, buffer overflows). This -separation keeps the NIF layer simple and the Elixir layer informative. - ---- - -## Scheduler Considerations - -### Why DirtyCpu? - -The BEAM scheduler is designed for fine-grained, cooperative concurrency. Each -BEAM process gets a budget of "reductions" (roughly, function calls). When a -process exhausts its budget, the scheduler preempts it and runs the next -process. - -A compression NIF call is not fine-grained. Compressing a 10 MB buffer with -Zstd at level 22 can take hundreds of milliseconds of continuous CPU time. -Running this on a normal scheduler thread blocks that thread from executing -other processes, causing: - -- Latency spikes for GenServer calls on the same scheduler -- Timeouts in process_link monitors -- Degraded cluster throughput - -The `schedule = "DirtyCpu"` annotation tells the BEAM to execute the NIF on a -dirty CPU scheduler: - -```rust -#[rustler::nif(schedule = "DirtyCpu")] -pub fn zstd_compress<'a>(env: Env<'a>, data: Binary, level: i32) -> Term<'a> { -``` - -### Dirty scheduler configuration - -The BEAM creates dirty CPU schedulers based on the `-SDio` flag (default: same -as online schedulers). A typical configuration: - -``` -+---- Normal schedulers (1 per core) ----+ -| Scheduler 1 | Scheduler 2 | ... | -| (Elixir processes, OTP tasks) | -+-----------------------------------------+ - -+---- Dirty CPU schedulers ----+ -| Dirty 1 | Dirty 2 | ... | -| (NIF calls) | -+-----------------------------+ - -+---- Dirty IO schedulers ----+ -| IO 1 | IO 2 | ... | -| (file, network I/O) | -+-----------------------------+ -``` - -All ExCodecs NIFs use `DirtyCpu` rather than `DirtyIo` because compression and -decompression are CPU-bound, not I/O-bound. - -### What happens without DirtyCpu? - -If a NIF is not annotated with `schedule = "DirtyCpu"`, Rustler defaults to -running it on a normal scheduler. For small inputs, this is fine. For inputs -larger than a few kilobytes, the NIF will block the scheduler thread for -potentially long periods, causing cascading latency in the BEAM. - -ExCodecs annotates every compression and decompression NIF with `DirtyCpu` -because the framework cannot predict input sizes at compile time. - ---- - -## Binary Handling - -### Input: Binary (immutable, reference-counted) - -Rustler provides the `Binary` type for reading BEAM binaries in Rust: - -```rust -pub fn zstd_compress<'a>(env: Env<'a>, data: Binary, level: i32) -> Term<'a> { -``` - -`Binary` is a zero-copy reference to the BEAM's binary data. When the BEAM -passes a binary to a NIF, it does not copy the data -- it provides a pointer -to the existing binary buffer. The `Binary` type ensures the binary's -reference count is maintained for the duration of the NIF call. - -Key properties: -- **Immutable**: The NIF cannot modify the input binary. -- **Zero-copy input**: No allocation or copying for the read path. -- **Slice access**: `data.as_slice()` returns a `&[u8]` view. - -### Output: NewBinary (mutable, allocated) - -Rustler's `NewBinary` type allocates a new BEAM binary of a known size: - -```rust -let mut output = NewBinary::new(env, compressed.len()); -output.as_mut_slice().copy_from_slice(&compressed); -(atoms::ok(), Binary::new(env, output.as_slice())).encode(env) -``` - -This is a two-step process: - -1. **Allocate**: `NewBinary::new(env, compressed.len())` allocates a BEAM - binary of exactly the right size. This avoids the overhead of Erlang's - binary append optimization (which over-allocates for growing binaries). - -2. **Copy**: `output.as_mut_slice().copy_from_slice(&compressed)` copies the - compressed data from the Rust `Vec` into the BEAM binary. - -3. **Encode**: `Binary::new(env, output.as_slice())` creates an immutable - `Binary` reference from the `NewBinary`, and `.encode(env)` converts it to - a BEAM term. - -### Binary allocation patterns across codecs - -Every codec follows the same allocation pattern: - -``` - +--------------+ +-----------------------+ +------------------+ - | BEAM Binary | | Rust Vec | | BEAM NewBinary | - | (input) | | (intermediate result) | | (output) | - +--------------+ +-----------------------+ +------------------+ - | | | - Binary::from compress/decompress NewBinary::new(env, len) - (zero-copy ref) produces Vec then copy_from_slice - | | | - v v v - data.as_slice() &compressed[..] output.as_slice() - | | | - +------- ALGORITHM ------+------- copy_from_slice ---+ -``` - -For all codecs except Blosc2, the intermediate `Vec` is the compressed or -decompressed data. For Blosc2, there is also a shuffle/unshuffle step in -between. - -### Memory management - -The Rust NIF has clear ownership boundaries: - -- **Input `Binary`**: Owned by the BEAM. The NIF borrows it for the call - duration. Rustler ensures the reference is valid. - -- **Intermediate `Vec`**: Owned by Rust. Allocated by the compression - crate, used for the algorithm, and then copied into the output binary. - -- **Output `NewBinary`**: Owned by the BEAM. Allocated in NIF context, - written to via `as_mut_slice()`, then converted to a BEAM term. The BEAM - takes ownership when the NIF returns. - -No Rust memory leaks are possible because all `Vec` values are dropped -when they go out of scope (Rust's ownership model). BEAM binaries are -garbage-collected normally. - ---- - -## Safety Considerations - -### No panics in NIF code - -A Rust panic inside a NIF will crash the entire BEAM process that called the -NIF. This is because panics in Rust unwind the stack, and the BEAM NIF -interface does not support unwinding across the NIF boundary. - -ExCodecs prevents panics through two mechanisms: - -1. **Comprehensive `match` on all Results**. Every compression/decompression - call returns a `Result`. The NIF code matches on `Ok` and `Err`, never - calling `.unwrap()` or `.expect()`. - -2. **Input validation in Elixir**. The Elixir codec modules validate all - options before calling the NIF. This means the NIF never receives - out-of-range values, null pointers, or empty binaries (except where - explicitly handled). - -```rust -// Safe: match on Result, never unwrap -match zstd::bulk::compress(data.as_slice(), level) { - Ok(compressed) => { - let mut output = NewBinary::new(env, compressed.len()); - output.as_mut_slice().copy_from_slice(&compressed); - (atoms::ok(), Binary::new(env, output.as_slice())).encode(env) - } - Err(_) => (atoms::error(), atoms::compression_failed()).encode(env), -} -``` - -### No unsafe code - -The ExCodecs Rust crate contains zero `unsafe` blocks. All memory operations -go through safe Rust abstractions: - -- `Binary::as_slice()` -- safe slice access -- `NewBinary::new()` -- safe allocation -- `as_mut_slice().copy_from_slice()` -- safe copy - -This is a deliberate design choice. If a codec crate requires `unsafe`, the -`unsafe` is confined to that crate's internals and is not exposed through the -ExCodecs NIF boundary. - -### Resource management - -All Rust resources follow the RAII (Resource Acquisition Is Initialization) -pattern. When a NIF call completes, all `Vec` buffers are dropped, all -`Binary` references are released, and the only surviving allocation is the -BEAM-owned `NewBinary` that was returned as the result. - -There are no Rustler `Resource` types in ExCodecs because the codec operations -are stateless -- each call is a pure function from binary input to binary -output. There is no persistent state to manage across NIF calls. - -### Integer overflow and clamping - -The NIF layer clamps integer inputs to valid ranges rather than returning -errors. This is consistent with the defense-in-depth philosophy: - -```rust -let level = level.clamp(1, 22); // Zstd levels -let block_size = block_size.clamp(1, 9); // Bzip2 block sizes -let clevel = clevel.clamp(0, 9); // Blosc2 compression levels -``` - -Clamping in the NIF is a safety net. The Elixir layer validates and rejects -invalid values with informative error messages; the NIF layer clamps as a -last resort. - -### Empty input handling - -The Blosc2 codec explicitly handles empty input by returning a valid Blosc2 -header with zero-length payload: - -```rust -if data.is_empty() { - let mut output = NewBinary::new(env, BLOSC_MIN_HEADER_LENGTH); - // Write header with nbytes=0 - return (atoms::ok(), Binary::new(env, output.as_slice())).encode(env); -} -``` - -This prevents division-by-zero or buffer-underflow errors in the shuffle and -compression logic. - ---- - -## Codec Implementation Patterns - -Every codec module in Rust follows the same structure. This section shows the -pattern, then highlights differences. - -### The standard pattern - -```rust -use rustler::{Binary, Encoder, Env, NewBinary, Term}; -use crate::atoms; - -pub fn version() -> String { - // Return the upstream crate version -} - -#[rustler::nif(schedule = "DirtyCpu")] -pub fn _compress<'a>(env: Env<'a>, data: Binary, ...) -> Term<'a> { - let /* validated/clamped params */ = ...; - - match ::compress(data.as_slice(), ...) { - Ok(compressed) => { - let mut output = NewBinary::new(env, compressed.len()); - output.as_mut_slice().copy_from_slice(&compressed); - (atoms::ok(), Binary::new(env, output.as_slice())).encode(env) - } - Err(_) => (atoms::error(), atoms::compression_failed()).encode(env), - } -} - -#[rustler::nif(schedule = "DirtyCpu")] -pub fn _decompress<'a>(env: Env<'a>, data: Binary) -> Term<'a> { - match ::decompress(data.as_slice()) { - Ok(decompressed) => { - let mut output = NewBinary::new(env, decompressed.len()); - output.as_mut_slice().copy_from_slice(&decompressed); - (atoms::ok(), Binary::new(env, output.as_slice())).encode(env) - } - Err(_) => (atoms::error(), atoms::decompression_failed()).encode(env), - } -} -``` - -### Zstd - -```rust -// Uses zstd::bulk::compress for one-shot compression -// Accepts level parameter (1-22, clamped in NIF) -// Uses zstd::decode_all for decompression - -#[rustler::nif(schedule = "DirtyCpu")] -pub fn zstd_compress<'a>(env: Env<'a>, data: Binary, level: i32) -> Term<'a> { - let level = level.clamp(1, 22); - match zstd::bulk::compress(data.as_slice(), level) { ... } -} -``` - -### LZ4 - -```rust -// Uses lz4_flex::compress / lz4_flex::decompress -// lz4_flex is a pure Rust implementation (no C FFI) -// Decompression requires a max_size hint; uses data.len() * 8 - -#[rustler::nif(schedule = "DirtyCpu")] -pub fn lz4_decompress<'a>(env: Env<'a>, data: Binary) -> Term<'a> { - let estimated_size = data.len() * 8; - let result = lz4_flex::decompress(data.as_slice(), estimated_size); - ... -} -``` - -### Snappy - -```rust -// Uses snap::raw::Encoder and snap::raw::Decoder (raw format, not framed) -// No configuration options -// Simplest codec in the set - -#[rustler::nif(schedule = "DirtyCpu")] -pub fn snappy_compress<'a>(env: Env<'a>, data: Binary) -> Term<'a> { - match snap::raw::Encoder::new().compress(data.as_slice()) { ... } -} -``` - -### Bzip2 - -```rust -// Uses bzip2::write::BzEncoder for compression (streaming API) -// Uses bzip2::read::BzDecoder for decompression (streaming API) -// Accepts block_size (1-9) and work_factor (0-250) - -#[rustler::nif(schedule = "DirtyCpu")] -pub fn bzip2_compress<'a>(env: Env<'a>, data: Binary, - block_size: u32, _work_factor: u32) -> Term<'a> { - let block_size = block_size.clamp(1, 9); - // Uses std::io::Write trait for compression - let mut writer = bzip2::write::BzEncoder::new( - &mut compressed, bzip2::Compression::new(block_size) - ); - writer.write_all(data.as_slice())?; - writer.finish()?; -} -``` - -### Blosc2 - -The Blosc2 codec is the most complex. Rather than binding to the C-Blosc2 -library via FFI, ExCodecs implements the Blosc2 format in pure Rust. - -This approach was chosen because: - -1. **No C dependency**: No system-level libblosc2 requirement. The NIF remains - pure Rust, compilable with `rustler_precompiled`. -2. **Format control**: The header format is well-defined and simple (16 bytes). - Parsing it in Rust is straightforward. -3. **Leverages existing Rust codecs**: The `internal_compress` and - `internal_decompress` functions delegate to the same `lz4_flex`, `snap`, - and `zstd` crates already used by the standalone codec NIFs. -4. **Shuffle implementation**: Byte shuffle and unshuffle are simple - transpose operations implemented in ~30 lines of Rust. - -The Blosc2 header format: - -``` -Offset Size Field -0 1 Magic byte (0x2c) -1 1 Version (2) -2 1 Flags (shuffle mode, uncompressed marker) -3 1 Reserved -4 4 nbytes (original size, little-endian u32) -8 4 cbytes (compressed size, little-endian u32) -12 1 cname (compressor code) -13 1 clevel (compression level) -14 1 shuffle (0=none, 1=byte, 2=bit) -15 1 typesize (element size in bytes) -16+ - Payload (compressed or uncompressed data) -``` - -When the compressed payload is larger than the original data, Blosc2 stores -the data uncompressed with a special flags byte (`0x01`). This is handled -explicitly in the compress path: - -```rust -if compressed.len() >= shuffled.len() && clevel > 0 { - // Store uncompressed with marker flag - flags = 0x01; - // Copy original data directly -} -``` - ---- - -## NIF Call Flow - -### Complete call flow: ExCodecs.encode(:zstd, data, level: 3) - -``` - 1. User calls ExCodecs.encode(:zstd, data, level: 3) - | - 2. ExCodecs.encode/3 validates codec name exists - | (is_atom and is_binary guards) - | - 3. CodecRegistry.lookup(:zstd) - | -> ETS lookup: {:zstd, {ExCodecs.Compression.Zstd, :compression, info}} - | -> {:ok, {ExCodecs.Compression.Zstd, :compression, info}} - | - 4. Check info.module != nil - | -> NIF is loaded, module exists - | - 5. validate_data(data) -> :ok - | - 6. ExCodecs.Compression.Zstd.encode(data, level: 3) - | - 7. Zstd.encode validates level (1-22) - | -> Keyword.get(opts, :level, 3) = 3 - | -> validate_level(3) = :ok - | - 8. ExCodecs.Native.zstd_compress(data, 3) - | -> Rustler dispatches to DirtyCpu scheduler - | - 9. Rust: zstd_codec::zstd_compress(env, data_binary, 3) - | -> level.clamp(1, 22) -> 3 - | -> zstd::bulk::compress(data.as_slice(), 3) - | -> Ok(compressed_vec) - | -10. Allocate BEAM binary: NewBinary::new(env, compressed_vec.len()) - | Copy: output.as_mut_slice().copy_from_slice(&compressed_vec) - | -11. Return: (atoms::ok(), Binary::new(env, output.as_slice())).encode(env) - | -> {:ok, compressed_binary} - | -12. Result propagates back through Zstd.encode -> ExCodecs.encode - | -13. {:ok, compressed_binary} returned to user -``` - -### Error call flow: ExCodecs.encode(:unknown, data) - -``` - 1. User calls ExCodecs.encode(:unknown, data, []) - | - 2. CodecRegistry.lookup(:unknown) - | -> ETS lookup returns [] - | -> {:error, :unsupported_codec} - | - 3. ExCodecs.encode matches {:error, :unsupported_codec} - | -> {:error, Error.new(:unsupported_codec, codec: :unknown)} - | - 4. {:error, %ExCodecs.Error{reason: :unsupported_codec, codec: :unknown}} - | returned to user -``` - -### Detailed NIF boundary flow - -``` - Elixir Process BEAM Runtime Rust NIF Thread - +------------------+ +------------------+ +------------------+ - | encode(:zstd, | | | | | - | data, level: 3) | | Dirty CPU | | zstd_compress() | - | | | | Scheduler | | | | - | v | | | | v | - | Zstd.encode() | | Process waits | schedule= | level = 3 | - | Native.zstd_ | -------- call --- | on dirty CPU | -- "DirtyCpu" ->| data.as_slice()| - | compress(data,3)| | thread to | | | | - | | | | complete | | v | - | ...waiting... | | | | zstd::bulk:: | - | | | | | | compress() | - | | | | | | | | - | | | <----- result --- | Result posted | <---- return ----| {:ok, binary} | - | v | | to process | | | - | {:ok, compressed} | +------------------+ +------------------+ - +------------------+ -``` - -Key observations: - -- The calling process blocks until the NIF completes. This is normal for - `DirtyCpu` NIFs -- the process is suspended and the scheduler thread is free - to run other processes. -- The dirty CPU scheduler thread runs the NIF to completion. It does not - participate in BEAM scheduling while the NIF is executing. -- The result is copied back to the BEAM process heap. The compressed binary is - allocated as a BEAM refc binary and the process receives a reference to it. \ No newline at end of file diff --git a/docs/spatial_formats.md b/docs/spatial_formats.md index e091e36..5fd852b 100644 --- a/docs/spatial_formats.md +++ b/docs/spatial_formats.md @@ -94,8 +94,9 @@ after `count` records are ignored. Truncation yields `:invalid_data` / `:truncated_input` as implemented. **Streaming:** `stream_decode` with `source: :file` reads the 16-byte header, -then each record via `IO.binread/2` (bounded memory). In-memory binaries still -materialize through `decode/2`. `stream_encode_to_file/3` requires `:schema` +then each record via `IO.binread/2` (bounded memory). In-memory binaries use +chunked Rust unpack when the spatial NIF is loaded; otherwise they materialize +through `decode/2`. `stream_encode_to_file/3` requires `:schema` (e.g. `schema: [:color]`) and patches the count after writing records. ## GSPL — `:gsplat` (version 1) @@ -127,8 +128,9 @@ with zeros on encode). currently informational; `sh_rest` count in the header is authoritative. **Streaming:** `stream_decode` with `source: :file` reads the 18-byte header, -then each record via `IO.binread/2` (bounded memory). In-memory binaries still -materialize through `decode/2`. `stream_encode_to_file/3` requires `:schema` +then each record via `IO.binread/2` (bounded memory). In-memory binaries use +chunked Rust unpack when the spatial NIF is loaded; otherwise they materialize +through `decode/2`. `stream_encode_to_file/3` requires `:schema` (e.g. `schema: []` or `schema: [sh_rest: 6]`) and patches the count after writing records. diff --git a/guides/understanding_bzip2.md b/guides/understanding_bzip2.md index 2d914d7..8a68aad 100644 --- a/guides/understanding_bzip2.md +++ b/guides/understanding_bzip2.md @@ -108,21 +108,11 @@ Larger block sizes allow the BWT to find more distant patterns, improving compre ## Work Factor -Bzip2's work factor parameter (0-250, default 30) controls the behavior when the BWT encounters "pathological" data that causes the sorting to take much longer than expected: - -- The default work factor of 30 means that if the sorting algorithm detects it is taking excessive time, it falls back to a simpler (but still correct) sorting method. -- Higher values make Bzip2 try harder before falling back. -- Lower values cause earlier fallback. - -In practice, the work factor rarely needs to be changed. Pathological data is uncommon, and the default of 30 handles nearly all real-world cases. - -```elixir -# Default work factor -{:ok, compressed} = ExCodecs.encode(:bzip2, data, block_size: 9, work_factor: 30) - -# Higher work factor for pathological data -{:ok, compressed} = ExCodecs.encode(:bzip2, data, block_size: 9, work_factor: 100) -``` +Bzip2's upstream `work_factor` parameter (0-250, default 30) governs fallback +behavior when the BWT encounters pathological data. **ExCodecs does not expose +it**: only `:block_size` is accepted on `encode/2`, and the NIF uses the +upstream default. Unknown options such as `work_factor:` are silently ignored, +so passing it has no effect. ## Bzip2 File Format diff --git a/guides/understanding_spatial_codecs.md b/guides/understanding_spatial_codecs.md index 8699e70..537d5ef 100644 --- a/guides/understanding_spatial_codecs.md +++ b/guides/understanding_spatial_codecs.md @@ -77,11 +77,22 @@ In ExCodecs that bag is a `GaussianCloud` of `Gaussian` structs. | Looks like geometry? | Sparse surface samples | Appearance-oriented volume of blobs | | ExCodecs types | `Point` / `PointCloud` | `Gaussian` / `GaussianCloud` | -## Pure Elixir - what “no NIFs” means +## Elixir-first, with optional Rust acceleration -Yes: **today there is no Rust/NIF accelerator for spatial codecs**. Encode and -decode run in Elixir. Compression codecs (`:zstd`, `:lz4`, …) still use Rust -NIFs; spatial formats do not. +Spatial encode/decode is **Elixir-first**: every format has a complete +pure-Elixir implementation, so the API works on any BEAM without a native +backend. Since **v0.2.3**, an optional Rust NIF accelerator is loaded when +the precompiled native library is available: + +- **Chunked Rust pack/unpack** for EXCP and GSPL (byte-compatible with the + Elixir path; verified by property tests). +- **mmap-backed file `stream_decode`** for EXCP/GSPL/PLY, and **binary PLY + body unpack**. +- **Chunked `stream_encode_to_file`** for EXCP/GSPL. + +Compression codecs (`:zstd`, `:lz4`, …) always use Rust NIFs; spatial formats +use the Elixir path unless acceleration is available. Pass `accel: false` to +force pure Elixir for a given call. A common pattern is: @@ -90,9 +101,18 @@ A common pattern is: {:ok, packed} = ExCodecs.encode(:zstd, ply) # NIF compresses the bytes ``` -Spatial work stays on the BEAM; you optionally compress the resulting binary -with the registry codecs. A future release may add an optional Rust backend -for spatial encode/decode while keeping the same Elixir API. +Spatial work stays on the BEAM (or the DirtyCpu NIF pool when accelerating); +you optionally compress the resulting binary with the registry codecs. + +### Streaming (v0.2.3+) + +EXCP / GSPL / PLY **files** decode record-by-record from disk (bounded +memory; mmap + Rust unpack when available). In-memory binaries use chunked +Rust unpack when the NIF is loaded, otherwise materialize through `decode/2`. +`stream_encode/2` still collects then encodes; use `encode_to_file/3` with an +explicit `:schema` for EXCP/GSPL incremental writes. Prefer explicit +`source: :file` or `source: :binary`. See `ExCodecs.Spatial.Stream` for +details. ## Domain types - what each one is for @@ -366,14 +386,9 @@ clouds. Magic bytes: `"GSPL"`. | “I heard splat / EXCP / GSPL” | splat = data kind; EXCP/GSPL = our binary containers | Wire layouts and schema rules are frozen in -[Spatial wire formats](../docs/spatial_formats.md). - -Stream helpers (v0.2.3+): EXCP / GSPL / PLY **files** decode record-by-record -from disk (bounded memory; mmap + Rust unpack when available). In-memory -binaries use chunked Rust unpack when the NIF is loaded, otherwise materialize. -`stream_encode/2` still collects then encodes; use `encode_to_file/3` with an -explicit `:schema` for EXCP/GSPL incremental writes. Prefer explicit -`source: :file` or `source: :binary`. +[Spatial wire formats](../docs/spatial_formats.md). Streaming behavior is +described in the [Elixir-first, with optional Rust acceleration](#elixir-first-with-optional-rust-acceleration) +section above and in `ExCodecs.Spatial.Stream`. ## Why a separate `ExCodecs.Spatial` API? diff --git a/lib/ex_codecs.ex b/lib/ex_codecs.ex index 3c8d685..870b2a3 100644 --- a/lib/ex_codecs.ex +++ b/lib/ex_codecs.ex @@ -267,7 +267,9 @@ defmodule ExCodecs do Lazily enumerates **spatial** primitives from a file path or binary. Delegates to `ExCodecs.Spatial.stream_decode/2`. **Not** used for registry - compression codecs. Today this materializes the payload then streams the list. + compression codecs. File sources stream record-by-record; in-memory binaries + materialize through `decode/2` (or use chunked Rust unpack when the spatial + NIF is loaded). ## Arguments diff --git a/lib/ex_codecs/application.ex b/lib/ex_codecs/application.ex index 4c4831b..9f9247e 100644 --- a/lib/ex_codecs/application.ex +++ b/lib/ex_codecs/application.ex @@ -102,15 +102,7 @@ defmodule ExCodecs.Application do nif_ok? = ExCodecs.Native.nif_loaded?() for {name, module, category, interface, metadata} <- codecs do - native? = - function_exported?(module, :__codec_info__, 0) and - match?(%{native?: true}, module.__codec_info__()) - - loadable? = - Code.ensure_loaded?(module) and function_exported?(module, :encode, 2) and - function_exported?(module, :decode, 2) - - if loadable? and (not native? or nif_ok?) do + if codec_available?(module, nif_ok?) do ExCodecs.CodecRegistry.register(name, module, category, interface, metadata) else ExCodecs.CodecRegistry.register_unavailable(name, category, interface) @@ -119,4 +111,18 @@ defmodule ExCodecs.Application do :ok end + + @doc false + def codec_available?(module, nif_ok?) do + loadable? = + Code.ensure_loaded?(module) and function_exported?(module, :encode, 2) and + function_exported?(module, :decode, 2) + + native? = + loadable? and + function_exported?(module, :__codec_info__, 0) and + match?(%{native?: true}, module.__codec_info__()) + + loadable? and (not native? or nif_ok?) + end end diff --git a/lib/ex_codecs/spatial.ex b/lib/ex_codecs/spatial.ex index 640ccce..2ad2a6a 100644 --- a/lib/ex_codecs/spatial.ex +++ b/lib/ex_codecs/spatial.ex @@ -46,11 +46,15 @@ defmodule ExCodecs.Spatial do ## Streaming note - EXCP (`:spatial_binary`) and GSPL (`:gsplat`) **file** sources stream - record-by-record from disk. PLY and in-memory binaries still materialize, - then enumerate. Prefer `source: :file` for large EXCP/GSPL paths, or - `source: :binary` for payloads. See `docs/spatial_formats.md` for `:auto` - path heuristics and wire-format layouts. + EXCP (`:spatial_binary`), GSPL (`:gsplat`), and PLY **file** sources + (`source: :file` or `:auto` path detection) stream record-by-record from + disk: only the header and one record are held in memory at a time. Binary + PLY uses a fixed stride; ASCII PLY reads lines. In-memory binaries still + materialize through `decode/2` before yielding, unless the spatial NIF is + loaded, in which case EXCP/GSPL use chunked Rust unpack. Prefer + `source: :file` for large paths, or `source: :binary` for payloads. See + `docs/spatial_formats.md` for `:auto` path heuristics and wire-format + layouts. """ alias ExCodecs.{CodecRegistry, Error} @@ -275,8 +279,11 @@ defmodule ExCodecs.Spatial do @doc """ Returns an enumerable over points or Gaussians decoded from a path or binary. - Despite the name, decoding currently materializes the complete payload and - decoded cloud before yielding elements. + File sources (`source: :file` or `:auto` path detection) stream + record-by-record: only the header and one record are held in memory at a + time. In-memory binaries materialize through `decode/2` before yielding, + unless the spatial NIF is loaded, in which case EXCP/GSPL use chunked Rust + unpack. See `ExCodecs.Spatial.Stream` for details. ## Arguments diff --git a/livebooks/01_introduction.livemd b/livebooks/01_introduction.livemd index 4c9f23d..a0cfab7 100644 --- a/livebooks/01_introduction.livemd +++ b/livebooks/01_introduction.livemd @@ -85,7 +85,7 @@ Let's see it in action. ```elixir original = String.duplicate(~S""" - Romeo and Juliet +Romeo and Juliet Excerpt from Act 2, Scene 2 JULIET @@ -119,31 +119,7 @@ Henceforth I never will be Romeo. JULIET What man art thou that thus bescreen'd in night So stumblest on my counsel? - """, 100) -``` - - - -``` -warning: outdented heredoc line. The contents inside the heredoc should be indented at the same level as the closing """. The following is forbidden: - - def text do - """ - contents - """ - end - -Instead make sure the contents are indented as much as the heredoc closing: - - def text do - """ - contents - """ - end - -The current heredoc line is indented too little -└─ work/ex_codecs/livebooks/01_introduction.livemd#cell:krsz4o6wgjua4fyq:3:31 - +""", 100) ``` @@ -219,23 +195,24 @@ Bzip2 block 9: 1311 bytes ### Blosc2 for Numerical Data ```elixir -# Blosc2 shines with typed binary data -data = :binary.copy(<<0, 0, 0, 0, 1, 1, 1, 1>>, 512) +# Float64 data: byte shuffle groups same-position bytes across elements +# (sign/exponent bytes cluster together), enabling much better compression. +floats = for i <- 1..2048, into: <<>>, do: <> -{:ok, plain} = ExCodecs.encode(:blosc2, data, typesize: 1, shuffle: :none) -{:ok, shuffled} = ExCodecs.encode(:blosc2, data, typesize: 1, shuffle: :byte) +{:ok, plain} = ExCodecs.encode(:blosc2, floats, typesize: 8, shuffle: :none) +{:ok, shuffled} = ExCodecs.encode(:blosc2, floats, typesize: 8, shuffle: :byte) -IO.puts("Original: #{byte_size(data)} bytes") -IO.puts("Blosc2 (no shuffle): #{byte_size(plain)} bytes") +IO.puts("Original: #{byte_size(floats)} bytes") +IO.puts("Blosc2 (no shuffle): #{byte_size(plain)} bytes") IO.puts("Blosc2 (byte shuffle): #{byte_size(shuffled)} bytes") ``` ``` -Original: 4096 bytes -Blosc2 (no shuffle): 73 bytes -Blosc2 (byte shuffle): 73 bytes +Original: 16384 bytes +Blosc2 (no shuffle): 8251 bytes +Blosc2 (byte shuffle): 675 bytes ``` diff --git a/livebooks/02_compression_fundamentals.livemd b/livebooks/02_compression_fundamentals.livemd index 34c01a0..b493607 100644 --- a/livebooks/02_compression_fundamentals.livemd +++ b/livebooks/02_compression_fundamentals.livemd @@ -380,13 +380,15 @@ On the BEAM these buffers live in DirtyCpu NIF memory, not the Erlang heap. ### Bzip2 block size — speed, ratio, and memory scale together -`:block_size` is 1..9. Each step is about **100 KiB** of compressor working -memory (and roughly half that on decompress). On short inputs the size win -from larger blocks is small; the interesting signal is encode/decode time -growing with block size. +`:block_size` is 1..9. Each step raises the block buffer by about **100 KiB** +(and roughly half that on decompress). To see block size actually bind, the +input below is ~1 MiB, so small block sizes split it into many blocks while +large ones use one or two. With that, larger blocks improve the ratio and +raise **encode** time (the BWT is superlinear per block); **decode** time is +roughly flat. Numbers are the mean of 5 runs after a warmup pass. ```elixir -data = String.duplicate("Hello, World! This is a compression test. ", 2000) +data = String.duplicate("Hello, World! This is a compression test. ", 25_000) IO.puts( String.pad_trailing("Block", 8) <> @@ -394,16 +396,30 @@ IO.puts( String.pad_trailing("Ratio%", 10) <> String.pad_trailing("Encode µs", 12) <> String.pad_trailing("Decode µs", 12) <> - "~Mem compress" + "Block buf" ) IO.puts(String.duplicate("-", 64)) for bs <- 1..9 do - {enc_us, {:ok, enc}} = :timer.tc(fn -> ExCodecs.encode(:bzip2, data, block_size: bs) end) - {dec_us, {:ok, ^data}} = :timer.tc(fn -> ExCodecs.decode(:bzip2, enc) end) + {:ok, enc} = ExCodecs.encode(:bzip2, data, block_size: bs) + ExCodecs.decode(:bzip2, enc) + + enc_times = + for _ <- 1..5 do + {us, {:ok, _}} = :timer.tc(fn -> ExCodecs.encode(:bzip2, data, block_size: bs) end) + us + end + + dec_times = + for _ <- 1..5 do + {us, {:ok, _}} = :timer.tc(fn -> ExCodecs.decode(:bzip2, enc) end) + us + end + + enc_us = div(Enum.sum(enc_times), 5) + dec_us = div(Enum.sum(dec_times), 5) ratio = Float.round(100 * byte_size(enc) / byte_size(data), 1) - # Rule of thumb from the Bzip2 guide: ~100 KiB × block_size while compressing. mem = "#{bs * 100} KiB" IO.puts( @@ -419,29 +435,33 @@ end IO.puts(""" Unlike Zstd, Bzip2 decode is also relatively slow — pick it for cold/archival -paths, not hot reads. Prefer a smaller block_size when concurrent compressions -would otherwise stack many megabytes of NIF memory. +paths, not hot reads. The "Block buf" column is just the block buffer +(≈ 100 KiB × block_size); total compressor memory adds fixed overhead on top. +Prefer a smaller block_size when concurrent compressions would otherwise stack +many megabytes of NIF memory. """) ``` ``` -Block Size Ratio% Encode µs Decode µs ~Mem compress +Block Size Ratio% Encode µs Decode µs Block buf ---------------------------------------------------------------- -1 158 0.2 20610 1363 100 KiB -2 158 0.2 15793 905 200 KiB -3 158 0.2 9738 599 300 KiB -4 158 0.2 9477 630 400 KiB -5 158 0.2 8405 539 500 KiB -6 158 0.2 8074 532 600 KiB -7 158 0.2 8427 504 700 KiB -8 158 0.2 7967 536 800 KiB -9 158 0.2 8143 499 900 KiB +1 1622 0.2 119100 5266 100 KiB +2 926 0.1 126151 5235 200 KiB +3 604 0.1 129753 5379 300 KiB +4 492 0.0 132152 5421 400 KiB +5 466 0.0 133521 5194 500 KiB +6 329 0.0 136100 5119 600 KiB +7 328 0.0 137070 5676 700 KiB +8 336 0.0 136531 5337 800 KiB +9 327 0.0 140045 5125 900 KiB Unlike Zstd, Bzip2 decode is also relatively slow — pick it for cold/archival -paths, not hot reads. Prefer a smaller block_size when concurrent compressions -would otherwise stack many megabytes of NIF memory. +paths, not hot reads. The "Block buf" column is just the block buffer +(≈ 100 KiB × block_size); total compressor memory adds fixed overhead on top. +Prefer a smaller block_size when concurrent compressions would otherwise stack +many megabytes of NIF memory. ``` diff --git a/livebooks/03_codec_comparison.livemd b/livebooks/03_codec_comparison.livemd index 67a5f48..bcb8598 100644 --- a/livebooks/03_codec_comparison.livemd +++ b/livebooks/03_codec_comparison.livemd @@ -226,36 +226,45 @@ VegaLite.new(width: 700, height: 350) ## Memory Usage -```elixir -memory_results = for codec <- codecs do - opts = - if codec == :blosc2, - do: [cname: :zstd, clevel: 5, shuffle: :byte, typesize: 8], - else: [] - {:ok, info} = ExCodecs.codec_info(codec) +NIF encode/decode allocates **off-heap refc binaries**, not process-heap +terms, so `Process.info(self(), :heap_size)` (which only sees the calling +process's heap) is the wrong metric for codec working memory. The cell below +measures process-heap growth across a full encode/decode of the 64 KiB +float array, with a garbage collect on each side: - mem_before = Process.info(self(), :heap_size) |> elem(1) +```elixir +for codec <- codecs do + opts = if codec == :blosc2, do: [cname: :zstd, clevel: 5, shuffle: :byte, typesize: 8], else: [] + :erlang.garbage_collect() + before_heap = Process.info(self(), :heap_size) |> elem(1) {:ok, enc} = ExCodecs.encode(codec, float_array, opts) {:ok, _dec} = ExCodecs.decode(codec, enc) - mem_after = Process.info(self(), :heap_size) |> elem(1) - - %{ - codec: inspect(codec), - category: info.category, - configurable: info.configurable?, - streaming: info.streaming?, - heap_growth_words: mem_after - mem_before - } + :erlang.garbage_collect() + after_heap = Process.info(self(), :heap_size) |> elem(1) + IO.puts(String.pad_trailing("#{codec}", 10) <> "heap growth: #{after_heap - before_heap} words") end - -Kino.DataTable.new(memory_results) ``` -```text -[%{configurable: false, category: :compression, codec: ":lz4", streaming: false, heap_growth_words: 0}, %{configurable: false, category: :compression, codec: ":snappy", streaming: false, heap_growth_words: 0}, %{configurable: true, category: :compression, codec: ":zstd", streaming: false, heap_growth_words: -8372}, %{configurable: true, category: :compression, codec: ":bzip2", streaming: false, heap_growth_words: 0}, %{configurable: true, category: :compression, codec: ":blosc2", streaming: false, heap_growth_words: 0}] ``` +lz4 heap growth: -988 words +snappy heap growth: 377 words +zstd heap growth: 0 words +bzip2 heap growth: 0 words +blosc2 heap growth: 0 words +``` + +The deltas are a few hundred words at most — Elixir-level term overhead +(`{:ok, enc}` tuples, bindings) — while the **64 KiB of binary data flowing +through each codec is off-heap refc memory** and does not appear here. That is +why naive `heap_size` snapshots (and even `:erlang.memory(:binary)`, which is +whole-VM and noisy) cannot yield a clean per-codec "memory usage" number. + +The real working-set bounds are set by options, not observed here: **decode** +is capped by `:max_output_size` (default 256 MiB, the decompression-bomb +guard), and **encode** working memory scales with the codec's block/window +size (e.g. bzip2 `block_size` × ~100 KiB; see livebook 02). ## Codec Profiles diff --git a/livebooks/04_building_storage_systems.livemd b/livebooks/04_building_storage_systems.livemd index 81ac5a8..4970dfc 100644 --- a/livebooks/04_building_storage_systems.livemd +++ b/livebooks/04_building_storage_systems.livemd @@ -318,19 +318,25 @@ Codec distribution: ```elixir defmodule ChunkedCompressor do @moduledoc """ - Compresses large data in chunks to limit memory usage. + Compresses large data in fixed-size chunks to limit peak memory usage. + + The container stores the chunk count, the original byte size, and an offset + table so individual chunks can be decompressed on demand (random access) + without re-decoding the whole blob. """ @default_chunk_size 64 * 1024 def compress(data, codec \\ :zstd, opts \\ [], chunk_size \\ @default_chunk_size) do - chunks = for <>, do: chunk + chunks = chunk_data(data, chunk_size) - compressed_chunks = Enum.map(chunks, fn chunk -> - {:ok, enc} = ExCodecs.encode(codec, chunk, opts) - enc - end) + compressed_chunks = + Enum.map(chunks, fn chunk -> + {:ok, enc} = ExCodecs.encode(codec, chunk, opts) + enc + end) - offsets = compressed_chunks + offsets = + compressed_chunks |> Enum.scan(0, fn chunk, acc -> acc + byte_size(chunk) end) |> Enum.drop(-1) |> List.insert_at(0, 0) @@ -346,27 +352,63 @@ defmodule ChunkedCompressor do end def decompress(blob, codec \\ :zstd) do - <> = blob + <> = blob offsets_size = num_chunks * 4 - <> = rest + <> = rest offsets_list = for <>, do: offset offsets_list = offsets_list ++ [byte_size(chunks_binary)] - chunks = Enum.chunk_every(offsets_list, 2, 1, :discard) + chunks = + Enum.chunk_every(offsets_list, 2, 1, :discard) |> Enum.map(fn [start_off, end_off] -> - size = end_off - start_off - <<_::binary-size(start_off), chunk::binary-size(size), _::binary>> = chunks_binary - chunk + binary_part(chunks_binary, start_off, end_off - start_off) end) - decompressed = Enum.map(chunks, fn chunk -> - {:ok, dec} = ExCodecs.decode(codec, chunk) - dec - end) + decompressed = + Enum.map(chunks, fn chunk -> + {:ok, dec} = ExCodecs.decode(codec, chunk) + dec + end) - {:ok, IO.iodata_to_binary(decompressed)} + result = IO.iodata_to_binary(decompressed) + + if byte_size(result) == original_size do + {:ok, result} + else + {:error, :size_mismatch} + end + end + + @doc """ + Decompresses a single chunk by 0-based index, using the offset table. + """ + def read_chunk(blob, index, codec \\ :zstd) do + <> = blob + + if index < 0 or index >= num_chunks do + {:error, :out_of_range} + else + offsets_size = num_chunks * 4 + <> = rest + + offsets_list = for <>, do: offset + offsets_list = offsets_list ++ [byte_size(chunks_binary)] + + [start_off, end_off] = Enum.at(Enum.chunk_every(offsets_list, 2, 1, :discard), index) + chunk = binary_part(chunks_binary, start_off, end_off - start_off) + + ExCodecs.decode(codec, chunk) + end + end + + defp chunk_data(data, chunk_size) when is_integer(chunk_size) and chunk_size > 0 do + case data do + <> -> [chunk | chunk_data(rest, chunk_size)] + <<>> -> [] + remainder when is_binary(remainder) -> [remainder] + end end end ``` @@ -383,19 +425,26 @@ large_data = String.duplicate("This is a chunk of data that repeats. ", 10000) {:ok, compressed} = ChunkedCompressor.compress(large_data, :zstd) {:ok, recovered} = ChunkedCompressor.decompress(compressed, :zstd) +<> = compressed +{:ok, first_chunk} = ChunkedCompressor.read_chunk(compressed, 0, :zstd) + IO.puts("Original: #{byte_size(large_data)} bytes") IO.puts("Compressed: #{byte_size(compressed)} bytes") IO.puts("Ratio: #{Float.round(100 * byte_size(compressed) / byte_size(large_data), 1)}%") +IO.puts("Chunks: #{num_chunks}") IO.puts("Round-trip: #{recovered == large_data}") +IO.puts("Random access chunk 0: #{byte_size(first_chunk)} bytes (matches: #{first_chunk == binary_part(large_data, 0, 64 * 1024)})") ``` ``` Original: 380000 bytes -Compressed: 322 bytes +Compressed: 384 bytes Ratio: 0.1% -Round-trip: false +Chunks: 6 +Round-trip: true +Random access chunk 0: 65536 bytes (matches: true) ``` diff --git a/livebooks/06_spatial_codecs.livemd b/livebooks/06_spatial_codecs.livemd index 4a61d8b..558f512 100644 --- a/livebooks/06_spatial_codecs.livemd +++ b/livebooks/06_spatial_codecs.livemd @@ -183,8 +183,9 @@ gaussian_cloud = `stream_decode/2` with `source: :file` reads one record/vertex at a time (bounded memory; mmap + Rust unpack when available). `stream_encode/2` still -collects the enumerable, then encodes once — use `encode_to_file/3` with an -explicit `:schema` for EXCP/GSPL incremental writes. +collects the enumerable, then encodes once — use +`ExCodecs.Spatial.Stream.encode_to_file/3` with an explicit `:schema` for +EXCP/GSPL incremental writes. ```elixir {:ok, streamed_payload} = @@ -198,6 +199,24 @@ streamed_points = length(streamed_points) ``` +```elixir +path = Path.join(System.tmp_dir!(), "ex_codecs_spatial_#{System.unique_integer([:positive])}.excp") + +:ok = + ExCodecs.Spatial.Stream.encode_to_file(points, path, + format: :spatial_binary, + schema: [] + ) + +file_points = + path + |> ExCodecs.Spatial.stream_decode(format: :spatial_binary, source: :file) + |> Enum.to_list() + +File.rm(path) +length(file_points) +``` + ## Category-Safe Dispatch The shared catalog does not overload the binary API with struct inputs. diff --git a/mix.exs b/mix.exs index 9ae0f1f..477f48b 100644 --- a/mix.exs +++ b/mix.exs @@ -92,7 +92,7 @@ defmodule ExCodecs.MixProject do "priv", "guides", "livebooks", - "docs", + "docs/spatial_formats.md", "checksum-*.exs", "mix.exs", "README.md", diff --git a/test/ex_codecs/application_test.exs b/test/ex_codecs/application_test.exs new file mode 100644 index 0000000..c33dd77 --- /dev/null +++ b/test/ex_codecs/application_test.exs @@ -0,0 +1,46 @@ +defmodule ExCodecs.ApplicationTest do + use ExUnit.Case, async: false + + alias ExCodecs.Application + + describe "codec_available?/2" do + test "a NIF-backed codec is unavailable when the NIF is not loaded" do + assert Application.codec_available?(NifBackedCodec, false) == false + end + + test "a NIF-backed codec is available when the NIF is loaded" do + assert Application.codec_available?(NifBackedCodec, true) == true + end + + test "a pure-Elixir codec is available regardless of NIF load state" do + assert Application.codec_available?(PureElixirCodec, false) == true + assert Application.codec_available?(PureElixirCodec, true) == true + end + + test "a module missing the codec callbacks is unavailable" do + assert Application.codec_available?(NotACodec, true) == false + end + end +end + +defmodule NifBackedCodec do + def __codec_info__ do + %{native?: true, streaming?: false, configurable?: false, version: "test"} + end + + def encode(_data, _opts), do: {:ok, <<>>} + def decode(_data, _opts), do: {:ok, <<>>} +end + +defmodule PureElixirCodec do + def __codec_info__ do + %{native?: false, streaming?: false, configurable?: false, version: "test"} + end + + def encode(_data, _opts), do: {:ok, <<>>} + def decode(_data, _opts), do: {:ok, <<>>} +end + +defmodule NotACodec do + def __codec_info__, do: %{native?: false, streaming?: false, configurable?: false} +end diff --git a/test/ex_codecs/spatial/accel_coverage_test.exs b/test/ex_codecs/spatial/accel_coverage_test.exs index 73e5a69..e27174b 100644 --- a/test/ex_codecs/spatial/accel_coverage_test.exs +++ b/test/ex_codecs/spatial/accel_coverage_test.exs @@ -5,6 +5,9 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do alias ExCodecs.Spatial.Codec.{Binary, Gsplat, PLY} alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} + unless ExCodecs.Spatial.Accel.available?(), + do: @moduletag(skip: "spatial Accel NIF not loaded") + defp tmp(ext) do Path.join( System.tmp_dir!(), @@ -12,14 +15,6 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do ) end - setup do - if Accel.available?() do - :ok - else - {:skip, "spatial Accel NIF not loaded"} - end - end - describe "Accel facade" do test "chunk_size, ply_type_tag, pack/unpack, mmap, and append_file" do assert Accel.chunk_size() == 4096 diff --git a/test/ex_codecs/spatial/accel_property_test.exs b/test/ex_codecs/spatial/accel_property_test.exs index 827fe27..7c5cdf8 100644 --- a/test/ex_codecs/spatial/accel_property_test.exs +++ b/test/ex_codecs/spatial/accel_property_test.exs @@ -8,13 +8,7 @@ defmodule ExCodecs.Spatial.AccelPropertyTest do @moduletag :accel - setup do - if Accel.available?() do - :ok - else - {:skip, "spatial Accel NIF not loaded"} - end - end + unless Accel.available?(), do: @moduletag(skip: "spatial Accel NIF not loaded") defp float32 do # Keep values in a range that survives f32 round-trip without NaN/Inf. From e3a2ef026f27fb1d9b38da4ce3094f1c5bcad16e Mon Sep 17 00:00:00 2001 From: thanos Date: Sat, 18 Jul 2026 15:14:22 -0400 Subject: [PATCH 5/6] =?UTF-8?q?All=20remaining=20items=20are=20now=20compl?= =?UTF-8?q?ete.=20Here's=20a=20summary=20of=20what=20was=20fixed=20in=20th?= =?UTF-8?q?is=20final=20pass:=20Code=20fixes=20-=20A-5:=20Fixed=20doc=20ex?= =?UTF-8?q?ample=20to=20use=20MyApp.CodecRegistry=20instead=20of=20the=20g?= =?UTF-8?q?lobal=20ExCodecs.CodecRegistry=20(avoids=20name=20collision)=20?= =?UTF-8?q?-=20CR-1:=20Added=20Code.ensure=5Floaded=3F/1=20guard=20in=20bu?= =?UTF-8?q?ild=5Fcodec=5Finfo/5=20before=20probing=20=5F=5Fcodec=5Finfo=5F?= =?UTF-8?q?=5F/0=20-=20CR-4:=20Documented=20the=20Agent's=20purpose=20(own?= =?UTF-8?q?s=20ETS=20table,=20state=20never=20read)=20-=20C-3:=20Documente?= =?UTF-8?q?d=20that=20validates=3F/1=20checks=20interface=20conformance,?= =?UTF-8?q?=20not=20data=20validation=20-=20C-4:=20Annotated=20Zstd=20type?= =?UTF-8?q?doc=20examples=20with=20NIF=20precondition=20-=20Z-1:=20Documen?= =?UTF-8?q?ted=20the=20Elixir/Rust=20level-clamp=20divergence=20in=20Zstd?= =?UTF-8?q?=20encode=20docs=20-=20Z-2:=20zstd=5Fversion/0=20now=20reads=20?= =?UTF-8?q?Native.codec=5Fversions()=20with=20a=20static=20fallback=20-=20?= =?UTF-8?q?ER-1:=20Unknown=20NIF=20error=20atoms=20now=20map=20to=20:inval?= =?UTF-8?q?id=5Fdata=20with=20the=20raw=20atom=20in=20details:=20-=20ER-2:?= =?UTF-8?q?=20Widened=20matches=3F/2=20spec=20to=20term()=20first=20argume?= =?UTF-8?q?nt=20-=20NF-2:=20Removed=20hardcoded=20268435456=20from=20the?= =?UTF-8?q?=20:output=5Flimit=5Fexceeded=20message=20-=20NF-3:=20Corrected?= =?UTF-8?q?=20safe=5Fcall/2=20spec=20to=20{:ok,=20binary()}=20-=20N-2:=20C?= =?UTF-8?q?ommented=20the=20nif=5Floaded=3F/0=20rescue=20clauses=20with=20?= =?UTF-8?q?concrete=20scenarios=20-=20N-3:=20Documented=20the=20base=5Furl?= =?UTF-8?q?=20=E2=86=94=20@source=5Furl=20coupling=20in=20native.ex=20-=20?= =?UTF-8?q?SP-5:=20resolve=5Fas/2=20now=20returns=20{:error,=20:invalid=5F?= =?UTF-8?q?options}=20instead=20of=20raising=20CaseClauseError=20-=20SP-6:?= =?UTF-8?q?=20Replaced=20all=20em-dashes=20and=20non-breaking=20hyphens=20?= =?UTF-8?q?with=20plain=20hyphens=20in=20spatial=20docs=20-=20SP-9:=20Mark?= =?UTF-8?q?ed=20Transform=20as=20"Reserved=20for=20future=20use"=20in=20it?= =?UTF-8?q?s=20moduledoc=20-=20SP-10:=20Documented=20the=20atom-key=20rest?= =?UTF-8?q?riction=20for=20literal=20%Point{}=20construction=20-=20SP-11:?= =?UTF-8?q?=20Added=20numeric=20guards=20to=20Bounds.new/2=20-=20SP-12:=20?= =?UTF-8?q?Added=20sh=20shape=20validation=20in=20Gaussian.new/2=20-=20SP-?= =?UTF-8?q?15:=20Restricted=20safe/1=20rescue=20to=20expected=20exception?= =?UTF-8?q?=20classes=20with=20warning=20logging=20-=20SP-16:=20Removed=20?= =?UTF-8?q?true=20as=20an=20accepted=20endianness=20atom=20-=20SP-19:=20EX?= =?UTF-8?q?CP=20decode=20now=20rejects=20unknown=20flag=20bits=20with=20:i?= =?UTF-8?q?nvalid=5Fdata=20-=20SP-20:=20Collapsed=203-pass=20Enum.any=3F?= =?UTF-8?q?=20+=20length/1=20into=20a=20single=20Enum.reduce=20-=20SP-23:?= =?UTF-8?q?=20Restricted=20PLY=20encode=20rescue=20to=20ArithmeticError/Fu?= =?UTF-8?q?nctionClauseError/ArgumentError=20-=20SP-24:=20Dropped=20redund?= =?UTF-8?q?ant=20metadata:=20attrs=20in=20PLY=20Gaussian=20decode=20-=20SP?= =?UTF-8?q?-25:=20Added=20validate=5Fcomments/1=20rejecting=20newlis=20and?= =?UTF-8?q?=20end=5Fheader=20in=20PLY=20comments=20-=20SP-27:=20Documented?= =?UTF-8?q?=20numeric-only=20attribute=20restriction=20in=20Point=20typedo?= =?UTF-8?q?c=20-=20R-5/R-6:=20Documented=20partial-record=20caller=20oblig?= =?UTF-8?q?ation=20and=20color/SH=20coercion=20in=20Accel=20moduledoc=20-?= =?UTF-8?q?=20R-8:=20Documented=20the=20N=C3=97256=20MiB=20amplification?= =?UTF-8?q?=20risk=20in=20Compression.decompress/3=20Documentation/test=20?= =?UTF-8?q?fixes=20-=20D-1:=20Verified=20repo=20slug=20thanos/codecs=20is?= =?UTF-8?q?=20consistent=20with=20git=20remote=20-=20D-2:=20Removed=20dupl?= =?UTF-8?q?icate=20checksum=20reminder=20from=200.2.0=20CHANGELOG=20-=20D-?= =?UTF-8?q?5:=20Added=20Keep=20a=20Changelog=20link=20references=20at=20bo?= =?UTF-8?q?ttom=20of=20CHANGELOG=20-=20D-10:=20Fixed=20codec=20chaining=20?= =?UTF-8?q?recommendation=20(spatial=E2=86=92zstd,=20not=20zstd=E2=86=92bl?= =?UTF-8?q?osc2)=20-=20T-13/T-17/T-18:=20Verified=20already=20fixed=20(exa?= =?UTF-8?q?ct=20pins,=20split=20tests,=20tmp=20cleanup)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- CHANGELOG.md | 14 +- config/test.exs | 4 - docs/architecture.md | 19 +- docs/spatial_formats.md | 10 +- guides/benchmarking_methodology.md | 70 +-- guides/choosing_compression_codec.md | 16 +- guides/codec_fundamentals.md | 2 +- guides/compression_fundamentals.md | 8 +- guides/native_architecture.md | 354 +++++----------- guides/runtime_codec_discovery.md | 2 +- guides/understanding_blosc2.md | 9 +- guides/understanding_bzip2.md | 15 +- guides/understanding_lz4.md | 12 +- guides/understanding_snappy.md | 64 ++- guides/understanding_zstd.md | 6 +- lib/ex_codecs/application.ex | 15 +- lib/ex_codecs/codec.ex | 82 +++- lib/ex_codecs/codec_registry.ex | 121 +++++- lib/ex_codecs/compression.ex | 10 +- lib/ex_codecs/compression/blosc2.ex | 24 +- lib/ex_codecs/compression/bzip2.ex | 18 +- lib/ex_codecs/compression/lz4.ex | 1 + lib/ex_codecs/compression/snappy.ex | 18 +- lib/ex_codecs/compression/zstd.ex | 20 +- lib/ex_codecs/error.ex | 2 +- lib/ex_codecs/native.ex | 5 + lib/ex_codecs/nif.ex | 78 +++- lib/ex_codecs/spatial.ex | 24 +- lib/ex_codecs/spatial/accel.ex | 80 +++- lib/ex_codecs/spatial/bounds.ex | 4 +- lib/ex_codecs/spatial/codec/binary.ex | 192 +++++---- lib/ex_codecs/spatial/codec/gsplat.ex | 156 ++++--- lib/ex_codecs/spatial/codec/ply.ex | 401 ++++++++++++------ lib/ex_codecs/spatial/codec/stream_source.ex | 51 +++ lib/ex_codecs/spatial/gaussian.ex | 20 +- lib/ex_codecs/spatial/point.ex | 11 + lib/ex_codecs/spatial/point_cloud.ex | 33 +- lib/ex_codecs/spatial/stream.ex | 40 +- lib/ex_codecs/spatial/transform.ex | 6 +- livebooks/01_introduction.livemd | 36 +- livebooks/03_codec_comparison.livemd | 62 +-- livebooks/04_building_storage_systems.livemd | 40 +- livebooks/05_zarr_style_workloads.livemd | 38 +- mix.exs | 24 +- native/ex_codecs_native/src/atoms.rs | 1 + native/ex_codecs_native/src/blosc2_codec.rs | 22 +- native/ex_codecs_native/src/spatial.rs | 12 +- native/ex_codecs_native/src/zstd_codec.rs | 7 +- test/ex_codecs/codec_registry_test.exs | 6 +- test/ex_codecs/codec_test.exs | 31 +- test/ex_codecs/compression/blosc2_test.exs | 8 + test/ex_codecs/compression/bzip2_test.exs | 17 +- test/ex_codecs/compression/lz4_test.exs | 10 +- test/ex_codecs/compression/snappy_test.exs | 10 +- .../compression/version_and_error_test.exs | 163 ------- .../compression/wire_interop_test.exs | 43 ++ test/ex_codecs/compression/zstd_test.exs | 22 +- test/ex_codecs/coverage_boost_test.exs | 23 +- test/ex_codecs/documentation_test.exs | 3 + test/ex_codecs/nif_test.exs | 2 +- test/ex_codecs/property_expansion_test.exs | 149 +++++++ test/ex_codecs/property_test.exs | 26 +- .../ex_codecs/spatial/accel_coverage_test.exs | 51 ++- test/ex_codecs/spatial/bounds_test.exs | 2 + test/ex_codecs/spatial/codec/binary_test.exs | 4 +- test/ex_codecs/spatial/coverage_test.exs | 46 +- test/ex_codecs/spatial/stream_test.exs | 2 + test/ex_codecs/spatial/types_test.exs | 6 + test/ex_codecs/spatial_test.exs | 10 +- test/fixtures/README.md | 24 ++ test/fixtures/bzip2/default.bin | Bin 0 -> 103 bytes test/fixtures/bzip2/default.src | 1 + test/fixtures/lz4/size_prepended.bin | Bin 0 -> 63 bytes test/fixtures/lz4/size_prepended.src | 1 + test/fixtures/snappy/raw.bin | Bin 0 -> 59 bytes test/fixtures/snappy/raw.src | 1 + test/fixtures/zstd/level3.bin | Bin 0 -> 69 bytes test/fixtures/zstd/level3.src | 1 + test/test_helper.exs | 2 +- 79 files changed, 1780 insertions(+), 1143 deletions(-) create mode 100644 lib/ex_codecs/spatial/codec/stream_source.ex delete mode 100644 test/ex_codecs/compression/version_and_error_test.exs create mode 100644 test/ex_codecs/compression/wire_interop_test.exs create mode 100644 test/ex_codecs/property_expansion_test.exs create mode 100644 test/fixtures/README.md create mode 100644 test/fixtures/bzip2/default.bin create mode 100644 test/fixtures/bzip2/default.src create mode 100644 test/fixtures/lz4/size_prepended.bin create mode 100644 test/fixtures/lz4/size_prepended.src create mode 100644 test/fixtures/snappy/raw.bin create mode 100644 test/fixtures/snappy/raw.src create mode 100644 test/fixtures/zstd/level3.bin create mode 100644 test/fixtures/zstd/level3.src diff --git a/CHANGELOG.md b/CHANGELOG.md index e19f046..b678ba9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -110,8 +110,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Notes -- Precompiled NIF checksums must be regenerated when publishing GitHub release - artifacts for `0.2.0`; until then local builds use `force_build` / source compile. - `native/ex_codecs_native/Cargo.lock` is committed for reproducible NIF builds. ## [0.1.1] - 2026-06-15 @@ -157,11 +155,17 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - Precompiled binaries for macOS (ARM64, x86_64), Linux (glibc, musl, ARM64), Windows (x86_64). - `ExCodecs.Compression` convenience module (`compress/3`, `decompress/3`). - Structured error handling with `%ExCodecs.Error{}`. -- 154 tests (unit + property-based with StreamData). -- 90%+ test coverage. - CI pipeline (test matrix, lint, coverage, docs). - Release workflow for Hex.pm publishing with precompiled NIFs. - Five livebooks: introduction, fundamentals, comparison, storage systems, Zarr workloads. - Eleven guides covering all public modules and API details. - Benchmarks via benchee. -- Credo and Dialyzer integration. \ No newline at end of file +- Credo and Dialyzer integration. + +[Unreleased]: https://github.com/thanos/codecs/compare/v0.2.3...HEAD +[0.2.3]: https://github.com/thanos/codecs/releases/tag/v0.2.3 +[0.2.2]: https://github.com/thanos/codecs/releases/tag/v0.2.2 +[0.2.1]: https://github.com/thanos/codecs/releases/tag/v0.2.1 +[0.2.0]: https://github.com/thanos/codecs/releases/tag/v0.2.0 +[0.1.1]: https://github.com/thanos/codecs/releases/tag/v0.1.1 +[0.1.0]: https://github.com/thanos/codecs/releases/tag/v0.1.0 \ No newline at end of file diff --git a/config/test.exs b/config/test.exs index cf741d0..df8c4c4 100644 --- a/config/test.exs +++ b/config/test.exs @@ -1,9 +1,5 @@ import Config -config :ex_codecs, - stream_data_runs: 100, - stream_data_max_size: 10_240 - config :excoveralls, output_dir: "cover/", coverage_options: [ diff --git a/docs/architecture.md b/docs/architecture.md index e5444e2..9e55d44 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -1,6 +1,6 @@ # ExCodecs Architecture -A production-quality, extensible BEAM-native codec framework for Elixir. +An extensible BEAM-native codec framework for Elixir. ## Table of Contents @@ -421,7 +421,7 @@ the need for users to have a Rust toolchain installed: # mix.exs defp deps do [ - {:rustler, "~> 0.36", runtime: false}, + {:rustler, "~> 0.36", optional: true}, {:rustler_precompiled, "~> 0.8"}, ] end @@ -748,7 +748,14 @@ category's preferred built-in order. | `:spatial_binary` | `ExCodecs.Spatial.Codec.Binary` | | `:gsplat` | `ExCodecs.Spatial.Codec.Gsplat` | -### Hashing +### Future categories (not implemented) + +The categories below are **planned but not shipped**. They are sketched here to +illustrate the extensibility model; the code examples are illustrative, not +runnable, and the registry entries shown do not exist today. `:irreversible_codec` +is not in the current `error_reason()` type. + +#### Hashing Hashing codecs encode data into fixed-length digests. Because hashing is one-way, `decode` will raise or return an error -- it is a symmetric API @@ -772,7 +779,7 @@ Planned hashing codecs: | BLAKE3 | Variable | Very fast, configurable output | | xxHash | 4/8 bytes | Non-cryptographic, very fast | -### Checksums +#### Checksums Checksum codecs produce small integrity tags. Unlike hashes, checksums may support a `verify` operation through encode/decode symmetry: @@ -792,7 +799,7 @@ Planned checksum codecs: | Adler32| 4 bytes | Faster than CRC32, less robust| | xxHash32| 4 bytes | Fast non-cryptographic | -### Binary Encodings +#### Binary Encodings Binary encoding codecs convert between binary formats and their text representations: @@ -816,7 +823,7 @@ Planned encoding codecs: | Base16 | 100% | Hex encoding | | Base58 | ~37% | Bitcoin-style, no ambiguous chars| -### Content Addressing +#### Content Addressing Content-addressed codecs combine hashing and encoding for use in content-addressable storage systems: diff --git a/docs/spatial_formats.md b/docs/spatial_formats.md index 5fd852b..04fdfa6 100644 --- a/docs/spatial_formats.md +++ b/docs/spatial_formats.md @@ -55,11 +55,13 @@ normal `{0,0,0}`). Mixed optional fields therefore round-trip with defaults fill - Pass `accel: false` on encode/decode/stream helpers to force pure Elixir. Prefer explicit `source: :file` or `source: :binary` when the argument is -ambiguous. File streams prefer mmap + DirtyCpu unpack when Accel is available. +ambiguous (`:binary` is the default). File streams prefer mmap + DirtyCpu +unpack when Accel is available. Do not truncate a mapped file while a stream +is live (SIGBUS risk). -`:auto` treats a binary as a path only when it looks path-like (under 4 KiB, no -`ply`/`EXCP`/`GSPL` magic, and contains `/` or `\` or ends with `.ply`/`.excp`/ -`.gspl`/`.bin`) **and** `File.regular?/1` is true. Edge cases: +`:auto` (opt-in) treats a binary as a path only when it looks path-like (under +4 KiB, no `ply`/`EXCP`/`GSPL` magic, and contains `/` or `\` or ends with +`.ply`/`.excp`/`.gspl`/`.bin`) **and** `File.regular?/1` is true. Edge cases: - A real file with no separator/extension is **not** auto-opened. - A short binary containing `/` that happens to be a regular file path **is** opened. diff --git a/guides/benchmarking_methodology.md b/guides/benchmarking_methodology.md index c56de42..72b7cd7 100644 --- a/guides/benchmarking_methodology.md +++ b/guides/benchmarking_methodology.md @@ -20,7 +20,7 @@ ExCodecs uses [Benchee](https://github.com/bencheeorg/benchee) for benchmarking. ```elixir # In mix.exs {:benchee, "~> 1.3", only: :bench}, -{:benchee_html, "~>> 1.0", only: :bench} +{:benchee_html, "~> 1.0", only: :bench} ``` ### Creating a Benchmark @@ -35,14 +35,12 @@ defp elixirc_paths(:bench), do: ["lib", "bench"] Create a benchmark file: ```elixir -# bench/compression_benchmark.exs -input_small = :crypto.strong_rand_bytes(1024) # 1 KB -input_medium = :crypto.strong_rand_bytes(1024 * 1024) # 1 MB -input_large = :crypto.strong_rand_bytes(10 * 1024 * 1024) # 10 MB - -# For realistic data, use your actual production data: -# input_json = File.read!("bench/data/sample.json") -# input_floats = File.read!("bench/data/float_array.bin") +# Illustrative sketch — prefer compressible fixtures over random bytes. +# The shipped suite is `bench/compression_bench.exs` (run via `mix benchmarks` +# or `MIX_ENV=bench mix run bench/run.exs`). +input_small = File.read!("bench/data/sample.json") # or your own fixture +input_medium = :binary.copy(input_small, 64) +input_large = :binary.copy(input_small, 640) Benchee.run( %{ @@ -53,9 +51,9 @@ Benchee.run( "blosc2_encode" => fn input -> ExCodecs.encode(:blosc2, input) end }, inputs: %{ - "1 KB" => input_small, - "1 MB" => input_medium, - "10 MB" => input_large + "small" => input_small, + "medium" => input_medium, + "large" => input_large }, time: 10, memory_time: 2, @@ -63,10 +61,12 @@ Benchee.run( ) ``` -Run it: +Run the shipped suite: ```bash -MIX_ENV=bench mix run bench/compression_benchmark.exs +mix benchmarks +# or +MIX_ENV=bench mix run bench/run.exs ``` ### Benchmarking with Options @@ -202,11 +202,12 @@ Reproducible benchmarks are essential for comparing results across different sys 1. **Pin the codec versions.** The versions of `zstd`, `lz4_flex`, `snap`, and `bzip2` crates in `Cargo.toml` determine the actual algorithm implementations. Record these versions. ```toml -# native/ex_codecs_native/Cargo.toml -zstd = "0.13" +# native/ex_codecs_native/Cargo.toml (record the versions you actually build) +structured-zstd = "0.0.48" lz4_flex = "0.11" snap = "1.1" -bzip2 = "0.4" +bzip2 = "0.6" +blosc2-pure-rs = "0.2.4" ``` 2. **Use fixed input data.** Generate test data deterministically or use a fixed seed: @@ -284,12 +285,14 @@ input = File.read!("bench/data/production_sample.json") Different data types compress differently. Always benchmark with data that matches your production workload: -| Data Type | Zstd L3 Ratio | LZ4 L1 Ratio | -|---------------|----------------|---------------| -| English text | 3.0:1 | 2.0:1 | -| JSON | 2.8:1 | 1.9:1 | -| Float64 array | 2.0:1 | 1.5:1 | -| Random bytes | 1.0:1 | 1.0:1 | +| Data Type | Zstd L3 Ratio | LZ4 Ratio | +|---------------|----------------|-----------| +| English text | 3.0:1 | 2.0:1 | +| JSON | 2.8:1 | 1.9:1 | +| Float64 array | 2.0:1 | 1.5:1 | +| Random bytes | 1.0:1 | 1.0:1 | + +(Illustrative order-of-magnitude ratios; benchmark your own data.) Results from one data type do not generalize to another. @@ -314,18 +317,16 @@ Decompression speed matters as much as compression speed for read-heavy workload ## Recommended Benchmark Suite -Create a comprehensive benchmark suite that covers these scenarios: +Create a benchmark suite that covers these scenarios: ### 1. Compression Throughput by Codec and Data Size ```elixir -# bench/throughput.exs +# Illustrative sketch — prefer compressible fixtures; random bytes only +# measure incompressible worst-case overhead. sizes = [ - {"1 KB", :crypto.strong_rand_bytes(1024)}, - {"10 KB", :crypto.strong_rand_bytes(10 * 1024)}, - {"100 KB", :crypto.strong_rand_bytes(100 * 1024)}, - {"1 MB", :crypto.strong_rand_bytes(1024 * 1024)}, - {"10 MB", :crypto.strong_rand_bytes(10 * 1024 * 1024)} + {"1 KB", File.read!("bench/data/sample.json")}, + {"1 MB", :binary.copy(File.read!("bench/data/sample.json"), 64)} ] codecs = [:zstd, :lz4, :snappy, :bzip2, :blosc2] @@ -425,7 +426,8 @@ end ### The Speed-Ratio Tradeoff -Plot compression ratio vs. speed for each codec and configuration: +Plot compression ratio vs. speed for each codec and configuration. +The ASCII plot below is **illustrative**, not measured on ExCodecs: ``` Ratio @@ -435,13 +437,13 @@ Ratio | * Zstd-5 3.0 | * Zstd-3 | * Zstd-1 -2.0 | * LZ4-1 * Snappy +2.0 | * LZ4 * Snappy | 1.0 +----------------------------------------- Speed - 500 MB/s 10 MB/s + fast slow ``` -Choose the point that meets your requirements. If you need >3:1 ratio and >100 MB/s, Zstd level 5 is a good choice. If you need >500 MB/s, LZ4 or Snappy are your options. +Choose the tradeoff that meets your requirements, then confirm with a Benchee run on your data. ### Context Matters diff --git a/guides/choosing_compression_codec.md b/guides/choosing_compression_codec.md index f5cc924..d4079ff 100644 --- a/guides/choosing_compression_codec.md +++ b/guides/choosing_compression_codec.md @@ -2,7 +2,7 @@ This guide helps you select the right compression codec for your use case. It includes a comparison table, decision criteria, and worked examples. -> ### Is the data spatial—or spatially representable? +> ### Is the data spatial, or spatially representable? > > Compression codecs operate on arbitrary binaries. If records represent > coordinates, points, normals, colors, oriented particles, or Gaussian splats, @@ -72,15 +72,15 @@ Answer the following questions to narrow your choice: ### 2. How fast does compression need to be? -- **Under 1 GB/s** -- LZ4 or Snappy. -- **Under 500 MB/s** -- Zstd level 1-3, Blosc2 with LZ4 inner codec. -- **No constraint** -- Zstd high levels (9-22) or Bzip2. +- **Need >= ~1 GB/s** - LZ4 or Snappy. +- **Need >= ~300-500 MB/s** - Zstd level 1-3, Blosc2 with LZ4 inner codec. +- **No constraint** - Zstd high levels (9-22) or Bzip2. ### 3. How fast does decompression need to be? -- **Multi-GB/s required** -- LZ4 or Snappy. -- **Fast, under 1 GB/s** -- Zstd any level, Blosc2. -- **No constraint** -- Any codec. Bzip2 decompression is typically 50-200 MB/s. +- **Need multi-GB/s** - LZ4 or Snappy. +- **Need ~1 GB/s class** - Zstd any level, Blosc2. +- **No constraint** - Any codec. Bzip2 decompression is typically much slower. ### 4. How much memory can you spare? @@ -204,7 +204,7 @@ A research pipeline archives float64 measurement arrays to cold storage. ) ``` -Rationale: The byte shuffle reorders bytes within each 8-byte float, grouping high-order bytes (often similar) together. Zstd then achieves ratio gains of 2-10x compared to compressing the raw array. +Rationale: The byte shuffle reorders bytes within each 8-byte float, grouping high-order bytes (often similar) together. On typed arrays this often yields about a 2-4x ratio gain versus compressing the raw array (illustrative; see the Blosc2 guide table). ### Example 4: Log File Archival diff --git a/guides/codec_fundamentals.md b/guides/codec_fundamentals.md index 0e85b4b..4b856c9 100644 --- a/guides/codec_fundamentals.md +++ b/guides/codec_fundamentals.md @@ -31,7 +31,7 @@ Key properties of this design: via `ExCodecs.Spatial`). - **Optionally configurable.** The `opts` keyword list allows codec-specific tuning (compression level, block size, shuffle mode) while keeping the call signature uniform. - **Explicit error handling.** Every operation returns `{:ok, result}` or `{:error, %ExCodecs.Error{}}`. There are no exceptions for normal failure modes. -- **Composable.** Because the input and output share the same type, codecs can be chained: encode with Zstd, then encode the result with Blosc2 if desired. +- **Composable.** Because the input and output share the same type, codecs can be chained. The sensible direction is to compress already-binary data with a fast codec (e.g. spatial `:ply` → `:zstd`); chaining two general-purpose compressors (Zstd-then-Blosc2) usually wastes time since the inner output is already high-entropy. - **Category modules** provide domain naming (`Compression.compress/3`) or non-binary shapes (`Spatial.encode/2`) without overloading the registry API. diff --git a/guides/compression_fundamentals.md b/guides/compression_fundamentals.md index 6c3af22..7eee9c3 100644 --- a/guides/compression_fundamentals.md +++ b/guides/compression_fundamentals.md @@ -132,7 +132,11 @@ Higher compression levels and larger window sizes consume more memory: ### Determinism -Some codecs (Snappy, LZ4) produce deterministic output for identical inputs. Others (Zstd at certain levels) may produce different compressed representations that decompress to the same data. +All ExCodecs backends produce deterministic output for a given input, fixed +options, and library version: the same `encode(:zstd, data, level: 3)` call +yields byte-identical output across calls. Output is **not** guaranteed across +library versions — a new structured-zstd release may produce a different (but +compatible) frame for the same input. ## When Compression Helps and Hurts @@ -162,4 +166,4 @@ Compression hurts when: 5. **Test with realistic data.** Synthetic benchmarks are misleading. Use production data or representative samples. -6. **Consider Blosc2 for typed arrays.** If your data is numerical (float arrays, integer matrices), Blosc2's shuffle filters can improve ratios by 2-10x compared to applying a general-purpose codec directly. \ No newline at end of file +6. **Consider Blosc2 for typed arrays.** If your data is numerical (float arrays, integer matrices), Blosc2's shuffle filters can improve ratios by about 2-4x compared to applying a general-purpose codec directly (illustrative; measure on your data). \ No newline at end of file diff --git a/guides/native_architecture.md b/guides/native_architecture.md index 934b64e..5499b50 100644 --- a/guides/native_architecture.md +++ b/guides/native_architecture.md @@ -1,299 +1,157 @@ # Native Architecture -ExCodecs uses Rust-based NIFs (Native Implemented Functions) via Rustler for high-performance compression operations. This guide explains how the native layer works, how it integrates with the BEAM, and how precompiled distribution ensures reliability. +ExCodecs implements compression and spatial unpack/pack in a Rust NIF crate +(`native/ex_codecs_native`), loaded through +[`RustlerPrecompiled`](https://github.com/philss/rustler_precompiled). This guide +describes the current crate layout, Dirty scheduler use, precompiled +distribution, and how Elixir wraps NIF results. -## Why Native Code? +## Why a NIF? -Compression algorithms perform computationally intensive operations on large binary data. Pure Elixir implementations would be orders of magnitude slower than C/Rust implementations because: +Compression and dense binary spatial decode are CPU-bound over large binaries. +Doing that work in Rust lets the library: -- **BEAM binary overhead**: Elixir binaries are managed by the BEAM garbage collector. Large binary processing in Elixir involves frequent allocations. -- **SIMD and cache efficiency**: Rust and C compilers can auto-vectorize tight loops over byte arrays, using CPU SIMD instructions (AVX2, NEON) that are not available in BEAM. -- **Memory layout**: Rust can work with contiguous byte slices (`&[u8]`) without copying, while BEAM binaries may require copying between BEAM and native memory. -- **Algorithm implementation quality**: Libraries like `zstd`, `lz4_flex`, and `bzip2` are highly optimized with years of performance tuning. +- Operate on contiguous `&[u8]` slices without BEAM binary traversal cost +- Use mature pure-Rust codec crates (no C compression toolchain at build time) +- Offload long work to Dirty CPU / Dirty IO schedulers so normal BEAM + schedulers stay responsive -The performance difference is significant: +There is **no** pure-Elixir Zstd/LZ4/Snappy/Bzip2/Blosc2 path. When the NIF is +absent, codecs register as unavailable and public APIs return +`%ExCodecs.Error{reason: :nif_not_loaded}`. -| Operation | Pure Elixir | Rust NIF | -|---------------|-------------|-----------| -| Zstd compress | ~5-20 MB/s | ~300+ MB/s| -| LZ4 compress | ~10-30 MB/s | ~500+ MB/s| -| Snappy compress| ~10-30 MB/s | ~500+ MB/s | +## Elixir entry point -For production workloads, the native implementation is not optional -- it is essential. - -## Rustler Integration - -ExCodecs uses [Rustler](https://github.com/rusterlium/rustler) to bridge between Elixir and Rust. Rustler provides: - -- **NIF generation**: Automatic generation of the NIF boilerplate from Rust function signatures. -- **Type conversion**: Safe conversion between Elixir terms and Rust types (binary to `&[u8]`, atoms, integers, etc.). -- **Error handling**: Rust `Result` types map to Elixir `{:ok, ...}` / `{:error, ...}` tuples. -- **Memory safety**: Rust's ownership model prevents buffer overflows, use-after-free, and other memory errors that are common in C NIFs. - -### The Native Module - -The Elixir side defines the NIF module: +`ExCodecs.Native` uses `RustlerPrecompiled` (not `use Rustler` alone): ```elixir -defmodule ExCodecs.Native do - use Rustler, - otp_app: :ex_codecs, - crate: :ex_codecs_native, - mode: :release - - def zstd_compress(_data, _level), do: :erlang.nif_error(:nif_not_loaded) - def zstd_decompress(_data), do: :erlang.nif_error(:nif_not_loaded) - def lz4_compress(_data, _level), do: :erlang.nif_error(:nif_not_loaded) - def lz4_decompress(_data), do: :erlang.nif_error(:nif_not_loaded) - # ... other NIFs - def codec_versions, do: :erlang.nif_error(:nif_not_loaded) - def nif_loaded?, do: not function_exported?(__MODULE__, :zstd_compress, 2) -end +use RustlerPrecompiled, + otp_app: :ex_codecs, + crate: :ex_codecs_native, + version: version, + base_url: "https://github.com/thanos/codecs/releases/download/v#{version}", + mode: :release, + nif_versions: ["2.17"], + targets: [ + "aarch64-apple-darwin", + "x86_64-apple-darwin", + "x86_64-unknown-linux-gnu", + "x86_64-unknown-linux-musl", + "aarch64-unknown-linux-gnu", + "aarch64-unknown-linux-musl", + "x86_64-pc-windows-msvc" + ] ``` -Each function has a fallback that returns `:erlang.nif_error(:nif_not_loaded)`. If the NIF library fails to load, calling any of these functions raises an error. The `nif_loaded?/0` function checks whether the NIF has been loaded by testing if `zstd_compress/2` is still the fallback. +Each exported function has an Elixir stub that raises +`:erlang.nif_error(:nif_not_loaded)` until the shared library loads. +`ExCodecs.Native.nif_loaded?/0` probes `codec_versions/0`. -### The Rust Crate +Public codecs never call `Native` stubs directly for error shaping: they go +through `ExCodecs.NIF.safe_call/2` / `wrap/2`, which map atoms and panics into +`%ExCodecs.Error{}`. Spatial codecs use `ExCodecs.Spatial.Accel` for the same +purpose. -The Rust crate lives in `native/ex_codecs_native/`: +## Rust crate layout ``` native/ex_codecs_native/ - Cargo.toml + Cargo.toml # rust-version = "1.94", pure-Rust deps src/ - lib.rs + lib.rs # rustler::init + codec_versions/0 atoms.rs - zstd_codec.rs - lz4_codec.rs - snappy_codec.rs - bzip2_codec.rs - blosc2_codec.rs -``` - -Each codec module implements the compression and decompression functions using the appropriate Rust library: - -```rust -// Simplified example of zstd_codec.rs -#[rustler::nif] -fn zstd_compress(data: Binary, level: i32) -> Result { - let compressed = zstd::encode_all(data.as_slice(), level) - .map_err(|e| Error::new(e.to_string()))?; - Ok(Binary::new(compressed)) -} - -#[rustler::nif] -fn zstd_decompress(data: Binary) -> Result { - let decompressed = zstd::decode_all(data.as_slice()) - .map_err(|e| Error::new(e.to_string()))?; - Ok(Binary::new(decompressed)) -} + zstd_codec.rs # structured-zstd + lz4_codec.rs # lz4_flex (size-prepended block) + snappy_codec.rs # snap (raw block) + bzip2_codec.rs # bzip2 0.6 (pure-Rust backend) + blosc2_codec.rs # blosc2-pure-rs + spatial.rs # EXCP / GSPL / PLY pack-unpack + mmap + util.rs ``` -The `Cargo.toml` specifies the Rust crate dependencies: +Current compression dependencies (no C `libzstd` / `libbzip2` / c-blosc2): ```toml -[dependencies] -rustler = "0.36" -zstd = "0.13" +structured-zstd = "0.0.48" lz4_flex = "0.11" snap = "1.1" -bzip2 = "0.4" +bzip2 = "0.6" +blosc2-pure-rs = { version = "0.2.4", features = ["zlib-rs"] } +memmap2 = "0.9" ``` -The release profile is optimized for performance: +Release profile: `opt-level = 3`, `lto = true`, `codegen-units = 1`, `strip = true`. -```toml -[profile.release] -opt-level = 3 -lto = true -codegen-units = 1 -strip = true -``` +### Version map -- **`opt-level = 3`**: Maximum optimization. -- **`lto = true`**: Link-Time Optimization, enabling cross-crate inlining. -- **`codegen-units = 1`**: Single codegen unit for better optimization. -- **`strip = true`**: Strip debug symbols for smaller binaries. +`codec_versions/0` returns a string map of backend labels, including +`"spatial" => "excp-gspl-ply-1"`. Elixir codec modules may hardcode library +crate versions in `__codec_info__/0`; treat the NIF map as a runtime probe, not +the only source of truth for HexDocs metadata. -## DirtyCpu Scheduling +## Dirty scheduling -Compression NIFs are CPU-intensive and can take significant time. The BEAM scheduler has two types of NIFs: +Dirty scheduling is **explicit** on each NIF (`#[rustler::nif(schedule = ...)]`). +Rustler does **not** auto-promote long NIFs. -- **Normal NIFs**: Must return quickly (under ~1 ms). If they take too long, they block the BEAM scheduler. -- **Dirty NIFs**: Long-running NIFs that are offloaded to dirty scheduler threads. +| Kind | Examples | +|------|----------| +| `DirtyCpu` | All compress/decompress; EXCP/GSPL/PLY pack and unpack (binary + mmap) | +| `DirtyIo` | `spatial_mmap_open`, `spatial_append_file` | -Rustler automatically marks NIFs as dirty when they are expected to take more than ~1 ms. Compression of data larger than a few kilobytes will take longer than this, so the NIFs in ExCodecs run on dirty CPU schedulers. +Configure dirty CPU count with the VM flag `+SDcpu N` (not `-SDio`). -### BEAM Scheduler Impact +Implications: -``` -BEAM Scheduler Threads -+-------------------------+ -| Scheduler 1 (normal) | Runs Elixir processes, short NIFs -| Scheduler 2 (normal) | -| Scheduler 3 (normal) | -| Scheduler 4 (normal) | -+-------------------------+ -| Dirty CPU 1 | Runs ExCodecs NIFs, other CPU NIFs -| Dirty CPU 2 | -+-------------------------+ -| Dirty IO 1 | Runs I/O NIFs -+-------------------------+ -``` +1. Normal schedulers stay free while compression runs on dirty threads. +2. Concurrent NIF work is limited by dirty scheduler count. +3. A started NIF cannot be cancelled. Wrapping in `Task.yield/2` + + `Task.shutdown/1` abandons the Elixir task; the dirty NIF may still finish + and hold a dirty scheduler until it returns. -The number of dirty CPU schedulers defaults to the number of CPU cores. You can configure this: +## Precompiled distribution -```elixir -# In vm.args or system flags -+SDcpu 4 # 4 dirty CPU schedulers -``` +CI attaches platform binaries to the GitHub release matching the package +version. At `mix deps.get` / app start, RustlerPrecompiled downloads the +matching artifact when present; otherwise it compiles from +`native/ex_codecs_native` (requires Rust **1.94+**). -### Implications for Production +`nif_versions: ["2.17"]` matches OTP 26+ NIF ABI 2.17. Supported OTP releases +for this package are those listed in `mix.exs` / README; precompiled binaries +target that NIF version. -1. **Dirty NIFs do not block normal schedulers.** Your Elixir processes continue running while compression happens in the background. +If load fails entirely, `ExCodecs.Application` registers codecs as unavailable +(`module: nil`) instead of crashing the VM. -2. **Concurrent compression is limited by dirty scheduler count.** If you dispatch more concurrent compression operations than dirty CPU schedulers, they will queue up. +## Error path -3. **Memory is allocated outside the BEAM heap.** NIF memory allocations use the system allocator, not the BEAM allocator. Large compression buffers do not trigger BEAM garbage collection but do consume system memory. - -4. **Monitors and timeouts are not possible for NIF calls.** Once a NIF starts, it runs to completion. You cannot set a timeout for a single NIF call. If you need timeouts, wrap the NIF call in a `Task` with `Task.yield/2`. - -```elixir -# Timeout pattern for NIF calls -task = Task.async(fn -> ExCodecs.encode(:bzip2, large_data) end) - -case Task.yield(task, 5000) || Task.shutdown(task) do - {:ok, result} -> result - nil -> {:error, :timeout} -end ``` - -## Precompiled Distribution - -ExCodecs uses [`rustler_precompiled`](https://github.com/philipatrojek/rustler_precompiled) to distribute pre-built NIF binaries, eliminating the need for a Rust compiler on the target machine. - -### How It Works - -1. During CI/release, NIF binaries are compiled for all target platforms. -2. The binaries are attached to the Hex package or fetched from a GitHub release. -3. At application startup, Rustler loads the precompiled binary for the current platform. -4. If no precompiled binary is available, Rustler falls back to compiling from source (requires Rust toolchain). - -The target platforms configured in `mix.exs`: - -```elixir -defp rustler_precompiled do - [ - targets: [ - "aarch64-apple-darwin", # macOS ARM64 - "x86_64-apple-darwin", # macOS x86_64 - "x86_64-unknown-linux-gnu", # Linux x86_64 (glibc) - "x86_64-unknown-linux-musl", # Linux x86_64 (musl/Alpine) - "aarch64-unknown-linux-gnu", # Linux ARM64 (glibc) - "aarch64-unknown-linux-musl", # Linux ARM64 (musl/Alpine) - "x86_64-pc-windows-msvc" # Windows x86_64 - ], - mode: :release, - nif_versions: ["2.17"] - ] -end +Rust Result::Err / atom + -> Rustler {:error, atom} or ErlangError + -> ExCodecs.NIF.wrap/2 or safe_call/2 + -> {:error, %ExCodecs.Error{}} ``` -### NIF Version Compatibility - -The `nif_versions: ["2.17"]` field specifies the BEAM NIF API version. OTP 27+ uses NIF version 2.17. This means: +There is no `Error.from_nif/2` helper. Panic-like `ErlangError` payloads are +logged and returned as `:compression_failed` with message +`"Native codec crashed"`. -- **OTP 27+**: Full compatibility. -- **OTP 26 and earlier**: May require compiling from source or using an older precompiled binary. - -Check your OTP version: - -```elixir -System.otp_release() -# => "27" -``` - -### Benefits of Precompiled NIFs - -- **No Rust toolchain required.** Users do not need to install `rustc`, `cargo`, or any Rust libraries. -- **Consistent builds.** Every platform gets the same optimized binary. -- **Fast installation.** `mix deps.get` downloads the precompiled binary; no compilation step. -- **Reduced CI time.** No Rust compilation in CI pipelines. - -### Fallback Behavior - -If the precompiled binary for the current platform is not available: - -1. Rustler attempts to compile the NIF from source in `native/ex_codecs_native/`. -2. This requires the Rust toolchain (`rustc`, `cargo`) and the Rust dependencies. -3. If compilation fails, the NIF is not loaded and all codecs are registered as unavailable. - -This fallback is transparent. The `ExCodecs.Native` module detects whether the NIF loaded successfully, and the application handles the unavailable case gracefully. - -## Codec Version Reporting - -The NIF provides a `codec_versions/0` function that returns the version of each underlying C/Rust library: - -```elixir -ExCodecs.Native.codec_versions() -# => %{ -# "zstd" => "1.5.6", -# "lz4" => "1.10.0", -# "snappy" => "1.1.10", -# "bzip2" => "0.4.4", -# "blosc2" => "2.14.0" -# } -``` - -Each codec module uses this to populate its `__codec_info__/0` metadata: - -```elixir -defp zstd_version do - case ExCodecs.Native.codec_versions() do - %{:zstd => v} -> v - _ -> "unknown" - end -rescue - _ -> "unknown" -end -``` - -The `rescue` clause handles the case where the NIF is not loaded, ensuring the module does not crash during compilation or when the NIF is unavailable. - -## Error Handling in Native Code - -When a Rust NIF encounters an error, it is converted to an Elixir error through several layers: - -1. **Rust side**: The compression library returns a `Result::Err`. Rustler converts this to an `{:error, reason}` term. -2. **Elixir side**: The codec module does not add additional error handling; the `{:error, reason}` tuple is returned directly. -3. **ExCodecs API**: The `encode/3` and `decode/3` functions wrap the error in an `%ExCodecs.Error{}` struct. - -```elixir -# Error wrapping in ExCodecs.Native -def from_nif({:error, reason}, codec) do - {:error, %ExCodecs.Error{ - reason: nif_error_to_atom(reason), - message: "NIF error in codec #{codec}: #{inspect(reason)}", - codec: codec, - details: reason - }} -end -``` +Decompression honors `:max_output_size` (default 256 MiB) to bound +amplification; see `ExCodecs.NIF` and `ExCodecs.Compression.decompress/3`. -Common NIF error scenarios: +## Spatial acceleration -- **Invalid data**: The decompression input is corrupt or not valid compressed data. -- **Memory allocation failure**: The system is out of memory. -- **Buffer overflow**: The decompressed size exceeds available memory. -- **Internal library error**: The underlying C/Rust library returned an unexpected error. +`ExCodecs.Spatial.Accel` is the Elixir ABI over the spatial NIFs: row tuple +shapes, chunked unpack (`{rows, next_offset}`), pack, and mmap resources. -All of these are caught and returned as `{:error, %ExCodecs.Error{}}` tuples, never as exceptions. +- Prefer `Accel.chunk_size/0` (4096) loops for large clouds. +- Do not truncate or replace a file while an mmap resource is live + (possible SIGBUS). Use `accel: false` / plain IO for untrusted paths. +- When `Accel.available?/0` is false, codecs fall back to pure Elixir decode. ## Summary -- ExCodecs uses Rustler NIFs for all compression operations, providing 10-100x speedup over pure Elixir. -- NIFs run on BEAM Dirty CPU schedulers, preventing them from blocking normal process scheduling. -- Precompiled binaries are distributed for 7 target platforms, requiring no Rust toolchain for installation. -- The registry enables graceful degradation when the NIF is unavailable. -- NIF errors are caught and returned as structured `ExCodecs.Error` tuples, never as exceptions. -- Release-mode compilation with LTO and `codegen-units = 1` ensures maximum performance. \ No newline at end of file +- Pure-Rust compression crates behind RustlerPrecompiled NIFs +- Explicit DirtyCpu / DirtyIo scheduling +- Structured errors via `ExCodecs.NIF`, never raw exceptions on the public API +- Optional spatial accel with Elixir fallback and mmap safety caveats diff --git a/guides/runtime_codec_discovery.md b/guides/runtime_codec_discovery.md index ae15256..635f9fd 100644 --- a/guides/runtime_codec_discovery.md +++ b/guides/runtime_codec_discovery.md @@ -344,6 +344,6 @@ In most deployments, all codecs are registered once at startup and never change. - The registry uses ETS for fast, lock-free lookups without process bottlenecks. - Codecs are registered at startup; unavailable ones (NIF not loaded) are marked but not removed. -- `supports?/1` and `available_codecs/0` enable robust runtime checks. +- `supports?/1` and `available_codecs/0` enable runtime checks. - Custom codecs can be registered by implementing the `ExCodecs.Codec` behaviour and calling `CodecRegistry.register/3`. - The distinction between "known" and "available" codecs enables graceful degradation when native libraries are missing. \ No newline at end of file diff --git a/guides/understanding_blosc2.md b/guides/understanding_blosc2.md index a3f0e36..1e26585 100644 --- a/guides/understanding_blosc2.md +++ b/guides/understanding_blosc2.md @@ -83,7 +83,7 @@ Blocks still improve cache utilization inside compression. ### Step 2: Shuffle Filters -The shuffle filter is Blosc2's most powerful feature for typed data. It reorders bytes within elements before compression: +The shuffle filter is Blosc2's main advantage for typed data. It reorders bytes within elements before compression: **Byte Shuffle** (`shuffle: :byte`): @@ -230,7 +230,7 @@ The `typesize` tells Blosc2 how many bytes each element occupies. It is critical Setting `typesize: 1` with `shuffle: :byte` is equivalent to `shuffle: :none` because there is no intra-element byte grouping. -Setting `typesize: 8` with `shuffle: :byte` on float64 arrays typically provides 2-10x ratio improvement over compressing the raw data directly. +Setting `typesize: 8` with `shuffle: :byte` on float64 arrays typically provides about a 2-4x ratio improvement over compressing the raw data directly (illustrative; measure on your arrays). **Important**: The data length should ideally be a multiple of `typesize` for the shuffle filter to work correctly. If the data length is not a multiple of `typesize`, Blosc2 will still work but the last partial element will not benefit from shuffling. @@ -283,7 +283,8 @@ or configurable native decompression threads. ## Ratio Improvement from Shuffle -The shuffle filter is Blosc2's key advantage for numerical data. Here are typical ratio improvements: +The shuffle filter is Blosc2's key advantage for numerical data. Illustrative +ratio improvements (measure on your arrays): | Data Type | Zstd Alone | Blosc2 + Zstd + Byte Shuffle | Improvement | |-------------------|------------|-------------------------------|-------------| @@ -292,7 +293,7 @@ The shuffle filter is Blosc2's key advantage for numerical data. Here are typica | Mixed JSON | 3.0:1 | 3.0:1 (shuffle: :none) | No change | | Random binary | 1.01:1 | 1.01:1 | No change | -For numerical data, the improvement from shuffle is dramatic and consistent. The more regular the data (arrays of identical types), the greater the benefit. +For numerical data, shuffle often helps substantially when the element layout is regular. ## Comparison with Other Codecs diff --git a/guides/understanding_bzip2.md b/guides/understanding_bzip2.md index 8a68aad..2447224 100644 --- a/guides/understanding_bzip2.md +++ b/guides/understanding_bzip2.md @@ -106,13 +106,8 @@ Larger block sizes allow the BWT to find more distant patterns, improving compre - **Block size 3-5**: Moderate ratio. Useful when memory is constrained. - **Block size 1-2**: Fastest Bzip2 compression. Only a small improvement over level 9 Zstd in ratio, but much slower. -## Work Factor - -Bzip2's upstream `work_factor` parameter (0-250, default 30) governs fallback -behavior when the BWT encounters pathological data. **ExCodecs does not expose -it**: only `:block_size` is accepted on `encode/2`, and the NIF uses the -upstream default. Unknown options such as `work_factor:` are silently ignored, -so passing it has no effect. +ExCodecs does not expose upstream Bzip2's `work_factor` parameter; only +`:block_size` is accepted on encode, and unknown keys are ignored. ## Bzip2 File Format @@ -136,7 +131,7 @@ A Bzip2 compressed stream has the following structure: - Optional run-length encoding **Stream Footer**: -- Magic bytes: `0x177245385090` (e ngth pi-related constant) +- Magic bytes: `0x177245385090` (sqrt(pi) as BCD digits) - CRC32 of the blocks - Padding bits @@ -231,6 +226,4 @@ Bzip2's advantage is raw compression ratio on text-heavy data. Zstd's advantage 4. **Do not use Bzip2 for small data.** The block header overhead and BWT setup cost make Bzip2 inefficient for payloads under 1 KB. -5. **Be aware of memory.** Each concurrent Bzip2 compression at block size 9 uses 9 MB. On the BEAM, plan for this when dispatching many concurrent compressions. - -6. **Leave the work factor at default.** Pathological data is rare. If compression seems to hang, increasing the work factor is a diagnostic step, not a production solution. \ No newline at end of file +5. **Be aware of memory.** Each concurrent Bzip2 compression at block size 9 uses 9 MB. On the BEAM, plan for this when dispatching many concurrent compressions. \ No newline at end of file diff --git a/guides/understanding_lz4.md b/guides/understanding_lz4.md index 55ad408..b1e13ae 100644 --- a/guides/understanding_lz4.md +++ b/guides/understanding_lz4.md @@ -110,7 +110,7 @@ LZ4 uses a 64 KB sliding window for match finding. This is fixed and not configu - Matches can reference data up to 64 KB back. - Patterns repeated at distances greater than 64 KB cannot be compressed. -- For data with long-range repetitions, Zstd (with its larger window) achieves significantly better ratios. +- For data with long-range repetitions, Zstd (typically a larger window than LZ4; not configurable in ExCodecs) achieves significantly better ratios. ## When LZ4 Excels @@ -125,7 +125,7 @@ LZ4 uses a 64 KB sliding window for match finding. This is fixed and not configu - **Cold storage.** Use Zstd or Bzip2 for better ratios when decompression speed is less important. - **Small, structured messages.** Consider Snappy for even simpler compression with no configuration. - **Numerical arrays.** Use Blosc2 with shuffle filters for better ratios on typed data. -- **Data with long-range patterns.** LZ4's 64 KB window misses patterns beyond 64 KB. Zstd's larger window captures them. +- **Data with long-range patterns.** LZ4's 64 KB window misses patterns beyond 64 KB. Zstd's typically larger window (not configurable in ExCodecs) can capture them. ## Comparison with Snappy @@ -133,9 +133,9 @@ LZ4 and Snappy occupy a similar niche (fast, low-ratio compression). Key differe | Property | LZ4 | Snappy | |-------------------|-------------------------------|-------------------------------| -| Compression Speed | ~500 MB/s (fixed fast profile)| ~500 MB/s | -| Decompression Speed | ~2 GB/s | ~1.5 GB/s | -| Compression Ratio | Moderate (better than Snappy)| Slightly lower | +| Compression Speed | Fast (fixed fast profile)| Fast | +| Decompression Speed | Very fast | Fast | +| Compression Ratio | Moderate (better than Snappy)| Slightly lower | | Configurable | No | No | | Format | LZ4 block format | Snappy framework format | | Deterministic | Yes | Yes | @@ -149,7 +149,7 @@ LZ4 generally offers better ratio at comparable speeds, while Snappy has a simpl | Compression Speed | Faster | Slower (at high levels)| | Decompression Speed| Faster | Fast | | Ratio | Moderate | High | -| Window Size | 64 KB (fixed) | Up to 8 MB+ (configurable) | +| Window Size | 64 KB (fixed) | Larger than LZ4 (not configurable in ExCodecs) | | Dictionary | No | Not exposed by ExCodecs | | Configurable | No | Levels 1-22 | diff --git a/guides/understanding_snappy.md b/guides/understanding_snappy.md index 8691a03..dee61fe 100644 --- a/guides/understanding_snappy.md +++ b/guides/understanding_snappy.md @@ -36,41 +36,36 @@ Snappy makes specific design choices that sacrifice ratio for speed: ## Snappy Format -The Snappy compressed format consists of a header and chunks of data. +ExCodecs uses the **raw Snappy block** format (not the separate Snappy framing format). A block is an uncompressed-length prefix followed by tagged elements. -### Stream Header +### Uncompressed Length Prefix ``` +-------------------+ -| Length (varint) | -- Decompressed length in bytes +| Length (varint) | -- Uncompressed length in bytes +-------------------+ ``` -The decompressed size is encoded as a variable-length integer at the start of the compressed output. This allows the decompressor to pre-allocate the exact output buffer size. +The uncompressed size is encoded as a variable-length integer at the start of the compressed block so the decompressor can size the output buffer. -### Chunks +### Elements -Each chunk is one of three types: +Each element starts with a tag byte. The lower two bits select the element type: -1. **Literal (tag byte: 0x00)**: Raw bytes that could not be compressed. -2. **Copy with 1-byte offset (tag byte: 0x01)**: Short copy, offset 0-63, length 4-7. -3. **Copy with 2-byte offset (tag byte: 0x02)**: Longer copy, offset 0-65535, length 4-1027. - -Each chunk is prefixed by a tag byte that encodes the chunk type and inline data: +| Tag (bits 1-0) | Type | +|----------------|------| +| `00` | Literal (raw bytes) | +| `01` | Copy with 1-byte offset (length 4-11, offset 0-2047) | +| `10` | Copy with 2-byte offset (length 1-64, offset 0-65535) | +| `11` | Copy with 4-byte offset (legacy; length 1-64) | ``` Tag byte: bits 7-2: length or offset info (depends on type) - bits 1-0: chunk type + bits 1-0: element type ``` -For literals, the length is encoded in the tag byte (up to 60 bytes inline, with extended length for larger literals). - -For copies: -- The offset is stored as a 1-byte or 2-byte little-endian value. -- The length minus the minimum match length (4) is stored inline in the tag byte or as an extra byte. - -The maximum offset is 64 KB. The maximum match length is 64 bytes for 1-byte offset copies and 64 KB for 2-byte offset copies. +For literals, short lengths fit in the tag; longer literals use a 1-4 byte length field. Copies encode offset and length as described in Google's Snappy format description. Typical compressor windows are up to 64 KB. ### Format Properties @@ -80,9 +75,13 @@ The maximum offset is 64 KB. The maximum match length is 64 bytes for 1-byte off ## Performance Characteristics +Speeds and ratios below are illustrative order-of-magnitude figures from +typical upstream C Snappy discussions; ExCodecs uses the pure-Rust `snap` +crate. Benchmark your own data. + ### Compression Speed -Snappy compresses at approximately 500 MB/s on modern hardware. Key factors: +Snappy is designed for very high compression throughput. Key factors: - The hash table is small and fits in L1/L2 cache. - No expensive operations (no sorting, no entropy coding, no optimal parsing). @@ -90,17 +89,15 @@ Snappy compresses at approximately 500 MB/s on modern hardware. Key factors: ### Decompression Speed -Snappy decompresses at approximately 1.5 GB/s. The decompressor: +Decompression is similarly speed-oriented. The decompressor: -- Reads the tag byte to determine the chunk type. +- Reads the tag byte to determine the element type. - For literals: copies bytes verbatim (essentially a memcpy). - For copies: reads the offset and length, then copies from the already-decompressed output. -This is nearly as fast as a memcpy of the compressed data, and the compressed data is smaller, so total time is often less than copying the uncompressed data. - ### Compression Ratio -Snappy's ratio is lower than LZ4, Zstd, and Bzip2. Typical ratios: +Snappy's ratio is lower than LZ4, Zstd, and Bzip2. Illustrative ratios: | Data Type | Snappy Ratio | LZ4 Ratio | Zstd Level 3 Ratio | |-------------|--------------|-----------|---------------------| @@ -113,7 +110,7 @@ Snappy's lower ratio is the expected tradeoff for its speed advantage. ## Using Snappy in ExCodecs -Snappy accepts no configuration options: +Snappy has no encode-time configuration knobs (`configurable?: false`): ```elixir # Compress @@ -123,15 +120,14 @@ Snappy accepts no configuration options: {:ok, original} = ExCodecs.decode(:snappy, compressed) ``` -The `opts` parameter is accepted but ignored: +Encode ignores option keywords. Decode accepts `:max_output_size` (same +decompression-bomb guard as other codecs): ```elixir -# Options are ignored -{:ok, compressed} = ExCodecs.encode(:snappy, data, []) +{:ok, compressed} = ExCodecs.encode(:snappy, data) +{:ok, original} = ExCodecs.decode(:snappy, compressed, max_output_size: 1_048_576) ``` -Snappy's codec metadata reflects this: - ```elixir {:ok, info} = ExCodecs.codec_info(:snappy) info.configurable? # => false @@ -167,10 +163,10 @@ Snappy and LZ4 serve a similar niche. The main differences: | Property | Snappy | LZ4 | |--------------------|---------------------|-----------------------------| -| Compression Speed | ~500 MB/s | ~500 MB/s (level 1) | -| Decompression Speed| ~1.5 GB/s | ~2 GB/s | +| Compression Speed | Fast | Fast (fixed fast profile) | +| Decompression Speed| Fast | Very fast | | Ratio | Slightly lower | Slightly higher | -| Configuration | None | 16 levels | +| Configuration | None | None (no levels in ExCodecs) | | Format complexity | Simple | Slightly more complex | | Deterministic | Yes | Yes | | Max offset | 64 KB | 64 KB | diff --git a/guides/understanding_zstd.md b/guides/understanding_zstd.md index 5138e7f..2d3dfde 100644 --- a/guides/understanding_zstd.md +++ b/guides/understanding_zstd.md @@ -167,9 +167,9 @@ On the BEAM, compression runs in a DirtyCpu NIF, so this memory is allocated out | Decompress Speed | Very Fast | Fast | Very Fast| Slow | | Dictionary Support| Yes | No | No | No | | Streaming | Yes | Yes | No | Yes | -| Configurable | 22 levels | 9 levels | 16 levels| 9 blocks | +| Configurable | 22 levels | 9 levels | No | 9 blocks | -Zstd is strictly superior to GZIP/Deflate on both ratio and speed. It is the recommended default for new systems. +Zstd is the recommended default for new systems: it usually offers a stronger speed/ratio tradeoff than GZIP/Deflate on modern workloads. Benchmark your own data rather than relying on published codec comparisons. ## Best Practices @@ -179,7 +179,7 @@ Zstd is strictly superior to GZIP/Deflate on both ratio and speed. It is the rec 3. **Use higher levels for write-once, read-many workloads.** The one-time compression cost is amortized over many fast decompressions. -4. **Reserve levels 15-22 for batch processing.** These levels can be 5-20x slower than level 3 for modest ratio improvements (typically 2-5% additional compression). +4. **Reserve levels 15-22 for batch processing.** Higher levels are often much slower than level 3 for modest extra ratio. Upstream C Zstd illustrations sometimes cite roughly 2-5% additional compression in that band; treat that as illustrative and benchmark your data (ExCodecs uses a pure-Rust backend). 5. **Watch memory at high levels.** Large payloads require room for input and output buffers (block API). diff --git a/lib/ex_codecs/application.ex b/lib/ex_codecs/application.ex index 9f9247e..4ff9be5 100644 --- a/lib/ex_codecs/application.ex +++ b/lib/ex_codecs/application.ex @@ -50,14 +50,17 @@ defmodule ExCodecs.Application do ## Callback implementation - An OTP application that embeds a similar registry can use the same shape: + An OTP application that embeds a similar registry can use the same shape. + Note: `ExCodecs.CodecRegistry` registers under a fixed global name, so the + example below is for a **separate** registry module, not a second instance + of the library's own. defmodule MyApp.Application do use Application @impl Application def start(_type, _args) do - children = [{ExCodecs.CodecRegistry, register: fn -> :ok end}] + children = [{MyApp.CodecRegistry, register: fn -> :ok end}] Supervisor.start_link(children, strategy: :one_for_one, name: MyApp.Supervisor @@ -74,13 +77,7 @@ defmodule ExCodecs.Application do opts = [strategy: :one_for_one, name: ExCodecs.Supervisor] - case Supervisor.start_link(children, opts) do - {:ok, pid} -> - {:ok, pid} - - error -> - error - end + Supervisor.start_link(children, opts) end @doc false diff --git a/lib/ex_codecs/codec.ex b/lib/ex_codecs/codec.ex index 1f9b42e..7f64bdc 100644 --- a/lib/ex_codecs/codec.ex +++ b/lib/ex_codecs/codec.ex @@ -15,11 +15,14 @@ defmodule ExCodecs.Codec do * `interface` (`:binary | :spatial`) — public API shape used by the entry * `module` (`module() | nil`) — implementation module, or `nil` when the codec is known but unavailable - * `native?` (`boolean()`) — whether the implementation uses a NIF - * `streaming?` (`boolean()`) — whether an incremental API is available - * `configurable?` (`boolean()`) — whether the codec accepts meaningful options + * `native?` (`boolean() | nil`) — whether the implementation uses a NIF + * `streaming?` (`boolean() | nil`) — whether an incremental API is available + * `configurable?` (`boolean() | nil`) — whether the codec accepts meaningful options * `version` (`String.t() | nil`) — backend version, if reported + Boolean capability fields and `version` may be `nil` on a bare `%ExCodecs.Codec{}` + until filled by `c:__codec_info__/0` and/or `ExCodecs.CodecRegistry.register/5`. + For example: iex> {:ok, %ExCodecs.Codec{} = codec} = ExCodecs.codec_info(:zstd) @@ -41,6 +44,7 @@ defmodule ExCodecs.Codec do # ... end + @impl true def __codec_info__ do %ExCodecs.Codec{ name: :zstd, @@ -64,6 +68,9 @@ defmodule ExCodecs.Codec do ## Example + Requires the NIF to be loaded (Zstd is NIF-backed). On a host without the + NIF, use a pure-Elixir codec or `ExCodecs.supports?(:zstd)` to gate. + iex> {:ok, result} = ExCodecs.Compression.Zstd.encode("codec input", []) iex> is_binary(result) true @@ -80,6 +87,8 @@ defmodule ExCodecs.Codec do ## Example + Requires the NIF to be loaded. + iex> {:ok, encoded} = ExCodecs.Compression.Zstd.encode("codec input", []) iex> ExCodecs.Compression.Zstd.decode(encoded, []) {:ok, "codec input"} @@ -193,6 +202,53 @@ defmodule ExCodecs.Codec do """ @callback decode(data :: binary(), opts :: keyword()) :: decode_result() + @doc """ + Returns catalog metadata for this codec module. + + Optional. When exported, `ExCodecs.CodecRegistry.register/5` reads capability + fields (`native?`, `streaming?`, `configurable?`, `version`) from the returned + struct unless the caller overrides them via the `metadata` keyword. Registry + arguments always supply `name`, `category`, `interface`, and `module` on the + stored entry. + + ## Arguments + + None. + + ## Returns + + An `%ExCodecs.Codec{}` whose capability fields describe this implementation. + Callers may leave `name`/`category`/`interface` unset when the registry fills + those from `register/5` arguments. + + ## Raises + + Implementations should not raise for a normal metadata return. Any exception + raised here propagates through registration. + + ## Implementation example + + defmodule ExCodecs.Compression.Zstd do + @behaviour ExCodecs.Codec + + @impl true + def __codec_info__ do + %ExCodecs.Codec{ + name: :zstd, + category: :compression, + module: __MODULE__, + native?: true, + streaming?: false, + configurable?: true, + version: "structured-zstd-0.0.48" + } + end + end + """ + @callback __codec_info__() :: t() + + @optional_callbacks __codec_info__: 0 + @typedoc """ Metadata for one shared codec-catalog entry. @@ -205,11 +261,16 @@ defmodule ExCodecs.Codec do * `interface` — `:binary` for the top-level registry API or `:spatial` for `ExCodecs.Spatial` * `module` — callback implementation, or `nil` for an unavailable codec - * `native?` — indicates NIF-backed operation - * `streaming?` — indicates incremental processing support - * `configurable?` — indicates codec-specific option support + * `native?` — indicates NIF-backed operation, or `nil` until known + * `streaming?` — indicates incremental processing support, or `nil` until known + * `configurable?` — indicates codec-specific option support, or `nil` until known * `version` — backend version string, or `nil` if unknown + Boolean capability fields default to `nil` on `%ExCodecs.Codec{}` and are + typically filled by `c:__codec_info__/0` and/or registry registration. After + a successful `register/5` or `register_unavailable/3` they are usually + concrete booleans; `nil` means unknown / not yet filled. + ## Example iex> {:ok, codec} = ExCodecs.codec_info(:zstd) @@ -222,9 +283,9 @@ defmodule ExCodecs.Codec do category: atom(), interface: :binary | :spatial, module: module() | nil, - native?: boolean(), - streaming?: boolean(), - configurable?: boolean(), + native?: boolean() | nil, + streaming?: boolean() | nil, + configurable?: boolean() | nil, version: String.t() | nil } @@ -240,6 +301,9 @@ defmodule ExCodecs.Codec do @doc """ Returns whether `module` exports `encode/2` and `decode/2`. + Despite the name, this checks **interface conformance** (the module + implements the codec callbacks), not runtime data validation. + ## Arguments * `module` (`module()`) — module to load and inspect diff --git a/lib/ex_codecs/codec_registry.ex b/lib/ex_codecs/codec_registry.ex index cfd6b0a..f440d66 100644 --- a/lib/ex_codecs/codec_registry.ex +++ b/lib/ex_codecs/codec_registry.ex @@ -7,6 +7,16 @@ defmodule ExCodecs.CodecRegistry do registry API, while spatial entries are dispatched through `ExCodecs.Spatial`. + ## Crash behavior + + The ETS table is owned by this Agent process. If the Agent **crashes**, the + table is destroyed and only the `:register` callback (built-ins) is replayed + on restart. **Runtime third-party registrations made via `register/3` or + `register/5` are lost** and must be re-registered by the caller. If the + Agent restarts while the table still exists (e.g. via a supervised + `:restart_for_type`), existing entries are preserved and only overwritten + by the callback. + ## Typical use iex> ExCodecs.available_codecs() @@ -18,6 +28,11 @@ defmodule ExCodecs.CodecRegistry do use Agent + # The Agent process exists solely to own the named public ETS table + # (`@table_name`). Its state is the table name itself and is never read + # after init; a GenServer would be equivalent. A future revision may move + # the catalog to `:persistent_term` for the read-mostly path. + @table_name :ex_codecs_registry @registry_name __MODULE__ @@ -73,7 +88,10 @@ defmodule ExCodecs.CodecRegistry do :ets.new(@table_name, [:set, :public, :named_table]) _tid -> - :ets.delete_all_objects(@table_name) + # Keep existing entries across registry restarts so third-party + # registrations survive and readers never see an empty catalog. + # Built-ins are overwritten by `register_fun` below. + :ok end register_fun.() @@ -162,7 +180,68 @@ defmodule ExCodecs.CodecRegistry do register(name, module, category, interface, []) end - @doc false + @doc """ + Registers a module with an explicit interface and optional metadata overrides. + + Builds the stored `%ExCodecs.Codec{}` from `name`, `module`, `category`, and + `interface`, then merges capability fields from the module's optional + `c:ExCodecs.Codec.__codec_info__/0` (when exported). Each key in `metadata` + overrides the corresponding field from `__codec_info__/0`. + + ## Arguments + + * `name` (`atom()`) — catalog key (always stored; not taken from metadata) + * `module` (`module()`) — module exporting `encode/2` and `decode/2` + * `category` (`atom()`) — grouping atom such as `:compression` or `:spatial` + * `interface` (`:binary | :spatial`) — public API shape for the entry + * `metadata` (`keyword()`) — optional overrides for capability fields: + + * `:native?` (`boolean()`) — NIF-backed implementation + * `:streaming?` (`boolean()`) — incremental API available + * `:configurable?` (`boolean()`) — meaningful codec-specific options + * `:version` (`String.t() | nil`) — backend version string + + Unknown keys are ignored. Omitted keys keep the value from + `__codec_info__/0`, or `nil` when that function is not exported. + + ## Returns + + * `:ok` — metadata was stored, replacing any entry with the same `name` + * `{:error, {:invalid_codec_module, module}}` — the module cannot be loaded + or does not export both codec callbacks + + ## Raises + + May raise `ArgumentError` if the registry ETS table does not exist. A codec's + optional `__codec_info__/0` is called during registration and any exception + it raises propagates. + + ## Persistence + + This entry is stored in an ETS table owned by the registry Agent. If the + Agent crashes, runtime registrations are **not replayed** — only the + `:register` callback supplied at `start_link/1` runs on restart. Re-register + third-party codecs from your own supervision tree if you need crash + survivability. + + ## Examples + + iex> :ok = ExCodecs.CodecRegistry.register( + ...> :documented_spatial_binary, + ...> ExCodecs.Spatial.Codec.Binary, + ...> :spatial, + ...> :spatial, + ...> native?: false, + ...> streaming?: false, + ...> configurable?: false, + ...> version: "EXCP 1" + ...> ) + iex> {:ok, info} = ExCodecs.CodecRegistry.codec_info(:documented_spatial_binary) + iex> {info.interface, info.version, info.native?} + {:spatial, "EXCP 1", false} + iex> ExCodecs.CodecRegistry.unregister(:documented_spatial_binary) + :ok + """ @spec register(atom(), module(), atom(), :binary | :spatial, keyword()) :: :ok | {:error, term()} def register(name, module, category, interface, metadata) @@ -209,7 +288,41 @@ defmodule ExCodecs.CodecRegistry do register_unavailable(name, category, :binary) end - @doc false + @doc """ + Registers a known codec as unavailable with an explicit interface. + + Stores `%ExCodecs.Codec{module: nil}` with capability fields set to `false` + and `version: nil`. Use when a codec name is part of the catalog but its + implementation module cannot be loaded (for example a NIF-backed codec when + the native library failed to load). + + ## Arguments + + * `name` (`atom()`) — known codec's registry key + * `category` (`atom()`) — grouping key, such as `:compression` or `:spatial` + * `interface` (`:binary | :spatial`) — public API shape for the entry + + ## Returns + + Always `:ok` after storing the unavailable entry. + + ## Raises + + May raise `ArgumentError` if the registry ETS table does not exist. + + ## Example + + iex> :ok = ExCodecs.CodecRegistry.register_unavailable( + ...> :example_spatial_optional, + ...> :spatial, + ...> :spatial + ...> ) + iex> {:ok, info} = ExCodecs.CodecRegistry.codec_info(:example_spatial_optional) + iex> {info.module, info.interface, ExCodecs.CodecRegistry.supports?(:example_spatial_optional)} + {nil, :spatial, false} + iex> ExCodecs.CodecRegistry.unregister(:example_spatial_optional) + :ok + """ @spec register_unavailable(atom(), atom(), :binary | :spatial) :: :ok def register_unavailable(name, category, interface) when interface in [:binary, :spatial] do @@ -483,7 +596,7 @@ defmodule ExCodecs.CodecRegistry do defp build_codec_info(name, module, category, interface, metadata) do info = - if function_exported?(module, :__codec_info__, 0) do + if Code.ensure_loaded?(module) and function_exported?(module, :__codec_info__, 0) do module.__codec_info__() else %ExCodecs.Codec{} diff --git a/lib/ex_codecs/compression.ex b/lib/ex_codecs/compression.ex index b941b83..6e9b49f 100644 --- a/lib/ex_codecs/compression.ex +++ b/lib/ex_codecs/compression.ex @@ -107,8 +107,12 @@ defmodule ExCodecs.Compression do * `codec` (`atom()`) — registered codec that produced the payload * `data` (`binary()`) — compressed bytes in that codec's wire format - * `opts` (`keyword()`) — codec-specific decode options; defaults to `[]` - and is currently ignored by the built-in compression codecs + * `opts` (`keyword()`) — decode options; defaults to `[]`. Built-in + compression codecs honor `:max_output_size` (positive integer bytes, + default 256 MiB). Do not raise the limit for untrusted inputs. + lz4/snappy/blosc2 allocate up to `max_output_size` before verifying + content, so N concurrent malicious requests can occupy N x 256 MiB; + tighten the limit for concurrent untrusted workloads. ## Returns @@ -118,6 +122,8 @@ defmodule ExCodecs.Compression do * `:unsupported_codec` when `codec` is not registered * `:codec_unavailable` when its implementation is unavailable * `:invalid_data` for invalid argument types or an unexpected NIF result + * `:invalid_options` when `:max_output_size` is not a positive integer + * `:output_limit_exceeded` when decompressed size would exceed the limit * `:decompression_failed` for corrupt, truncated, or mismatched input * `:nif_not_loaded` when the native library is unavailable diff --git a/lib/ex_codecs/compression/blosc2.ex b/lib/ex_codecs/compression/blosc2.ex index 1bf77c4..61e8761 100644 --- a/lib/ex_codecs/compression/blosc2.ex +++ b/lib/ex_codecs/compression/blosc2.ex @@ -19,8 +19,8 @@ defmodule ExCodecs.Compression.Blosc2 do | B2ND (N-D arrays) | No | For large data, split buffers yourself and store multiple chunks. Usability - for “compress my array slice in Elixir” is full; for “Blosc2 as a database - file format”, use python/C Blosc2. + for "compress my array slice in Elixir" is full; for "Blosc2 as a database + file format", use python/C Blosc2. ## Options @@ -100,6 +100,7 @@ defmodule ExCodecs.Compression.Blosc2 do version: "c-blosc2-chunk/pure-rust" } """ + @impl true def __codec_info__ do %ExCodecs.Codec{ name: :blosc2, @@ -202,15 +203,19 @@ defmodule ExCodecs.Compression.Blosc2 do * `data` (`binary()`) — one chunk produced by this codec or by C/python-blosc2 single-buffer `compress` - * `opts` (`term()`) — ignored by this direct function; callers using the - codec behaviour or registry API should pass the keyword list `[]` + * `opts` (`keyword()`) — optional `:max_output_size` (positive integer + bytes, default 256 MiB) ## Returns * `{:ok, decompressed :: binary()}` on success * `{:error, %ExCodecs.Error{reason: :invalid_data}}` when `data` is not a - binary, the chunk header is shorter than 16 bytes, the declared output - exceeds 1 GiB, or the NIF raises an argument error + binary, `opts` is not a list, the chunk header is shorter than 16 bytes, + or the NIF raises an argument error + * `{:error, %ExCodecs.Error{reason: :invalid_options}}` when + `:max_output_size` is not a positive integer + * `{:error, %ExCodecs.Error{reason: :output_limit_exceeded}}` when the + declared chunk nbytes exceeds `:max_output_size` or the 1 GiB hard cap * `{:error, %ExCodecs.Error{reason: :decompression_failed}}` when the chunk is corrupt, truncated, unsupported, or is a Blosc2 container rather than a single chunk @@ -219,10 +224,9 @@ defmodule ExCodecs.Compression.Blosc2 do ## Raises / Exceptions - Data guard failures and `ErlangError`/`ArgumentError` exceptions from the NIF - call are converted to error tuples. Because `opts` is ignored, this direct - function also accepts non-list option terms. Unexpected exception classes - may propagate. + Guard failures and `ErlangError`/`ArgumentError` exceptions from the NIF + call are converted to error tuples. Unexpected exception classes may + propagate. ## Examples diff --git a/lib/ex_codecs/compression/bzip2.ex b/lib/ex_codecs/compression/bzip2.ex index 9ff37b9..8fdacfd 100644 --- a/lib/ex_codecs/compression/bzip2.ex +++ b/lib/ex_codecs/compression/bzip2.ex @@ -63,6 +63,7 @@ defmodule ExCodecs.Compression.Bzip2 do version: "bzip2-0.6/libbz2-rs" } """ + @impl true def __codec_info__ do %ExCodecs.Codec{ name: :bzip2, @@ -136,14 +137,18 @@ defmodule ExCodecs.Compression.Bzip2 do ## Arguments * `data` (`binary()`) — a complete Bzip2 stream - * `opts` (`term()`) — ignored by this direct function; callers using the - codec behaviour or registry API should pass the keyword list `[]` + * `opts` (`keyword()`) — optional `:max_output_size` (positive integer + bytes, default 256 MiB) ## Returns * `{:ok, decompressed :: binary()}` on success * `{:error, %ExCodecs.Error{reason: :invalid_data}}` when `data` is not a - binary or the NIF raises an argument error + binary, `opts` is not a list, or the NIF raises an argument error + * `{:error, %ExCodecs.Error{reason: :invalid_options}}` when + `:max_output_size` is not a positive integer + * `{:error, %ExCodecs.Error{reason: :output_limit_exceeded}}` when the + decompressed size would exceed `:max_output_size` * `{:error, %ExCodecs.Error{reason: :decompression_failed}}` when `data` is corrupt, truncated, or not a Bzip2 stream * `{:error, %ExCodecs.Error{reason: :nif_not_loaded}}` when the native @@ -151,10 +156,9 @@ defmodule ExCodecs.Compression.Bzip2 do ## Raises / Exceptions - Data guard failures and `ErlangError`/`ArgumentError` exceptions from the NIF - call are converted to error tuples. Because `opts` is ignored, this direct - function also accepts non-list option terms. Unexpected exception classes - may propagate. + Guard failures and `ErlangError`/`ArgumentError` exceptions from the NIF + call are converted to error tuples. Unexpected exception classes may + propagate. ## Examples diff --git a/lib/ex_codecs/compression/lz4.ex b/lib/ex_codecs/compression/lz4.ex index a3fceac..659e34b 100644 --- a/lib/ex_codecs/compression/lz4.ex +++ b/lib/ex_codecs/compression/lz4.ex @@ -59,6 +59,7 @@ defmodule ExCodecs.Compression.Lz4 do version: "lz4_flex-0.11" } """ + @impl true def __codec_info__ do %ExCodecs.Codec{ name: :lz4, diff --git a/lib/ex_codecs/compression/snappy.ex b/lib/ex_codecs/compression/snappy.ex index d745c46..12e7ab1 100644 --- a/lib/ex_codecs/compression/snappy.ex +++ b/lib/ex_codecs/compression/snappy.ex @@ -60,6 +60,7 @@ defmodule ExCodecs.Compression.Snappy do version: "snap-1.1" } """ + @impl true def __codec_info__ do %ExCodecs.Codec{ name: :snappy, @@ -123,14 +124,18 @@ defmodule ExCodecs.Compression.Snappy do * `data` (`binary()`) — a standalone Snappy block, normally produced by `encode/2` - * `opts` (`term()`) — ignored by this direct function; callers using the - codec behaviour or registry API should pass the keyword list `[]` + * `opts` (`keyword()`) — optional `:max_output_size` (positive integer + bytes, default 256 MiB) ## Returns * `{:ok, decompressed :: binary()}` on success * `{:error, %ExCodecs.Error{reason: :invalid_data}}` when `data` is not a - binary or the NIF raises an argument error + binary, `opts` is not a list, or the NIF raises an argument error + * `{:error, %ExCodecs.Error{reason: :invalid_options}}` when + `:max_output_size` is not a positive integer + * `{:error, %ExCodecs.Error{reason: :output_limit_exceeded}}` when the + decompressed size would exceed `:max_output_size` * `{:error, %ExCodecs.Error{reason: :decompression_failed}}` when `data` is corrupt, truncated, or not a Snappy block * `{:error, %ExCodecs.Error{reason: :nif_not_loaded}}` when the native @@ -138,10 +143,9 @@ defmodule ExCodecs.Compression.Snappy do ## Raises / Exceptions - Data guard failures and `ErlangError`/`ArgumentError` exceptions from the NIF - call are converted to error tuples. Because `opts` is ignored, this direct - function also accepts non-list option terms. Unexpected exception classes - may propagate. + Guard failures and `ErlangError`/`ArgumentError` exceptions from the NIF + call are converted to error tuples. Unexpected exception classes may + propagate. ## Examples diff --git a/lib/ex_codecs/compression/zstd.ex b/lib/ex_codecs/compression/zstd.ex index 8793941..57dadb5 100644 --- a/lib/ex_codecs/compression/zstd.ex +++ b/lib/ex_codecs/compression/zstd.ex @@ -74,6 +74,7 @@ defmodule ExCodecs.Compression.Zstd do version: "structured-zstd-0.0.48" } """ + @impl true def __codec_info__ do %ExCodecs.Codec{ name: :zstd, @@ -86,7 +87,18 @@ defmodule ExCodecs.Compression.Zstd do } end - defp zstd_version, do: "structured-zstd-0.0.48" + defp zstd_version do + case ExCodecs.Native.nif_loaded?() do + true -> + versions = ExCodecs.Native.codec_versions() + Map.get(versions, "zstd", "structured-zstd-0.0.48") + + false -> + "structured-zstd-0.0.48" + end + rescue + _ -> "structured-zstd-0.0.48" + end @doc """ Compresses a binary with pure-Rust Zstd. @@ -115,6 +127,12 @@ defmodule ExCodecs.Compression.Zstd do exceptions from the NIF call are converted to error tuples. Unexpected exception classes may propagate. + ## Notes + + The Elixir layer rejects out-of-range `:level` with `:invalid_options`, + while the Rust NIF silently clamps. Direct `ExCodecs.Native` callers bypass + this validation. + ## Examples iex> payload = :binary.copy("telemetry,", 20) diff --git a/lib/ex_codecs/error.ex b/lib/ex_codecs/error.ex index 8be5cd2..c3b19ab 100644 --- a/lib/ex_codecs/error.ex +++ b/lib/ex_codecs/error.ex @@ -233,7 +233,7 @@ defmodule ExCodecs.Error do iex> ExCodecs.Error.matches?({:error, error}, :invalid_data) false """ - @spec matches?({:error, t()}, error_reason()) :: boolean() + @spec matches?(term(), error_reason()) :: boolean() def matches?({:error, %__MODULE__{reason: reason}}, reason), do: true def matches?(_, _), do: false end diff --git a/lib/ex_codecs/native.ex b/lib/ex_codecs/native.ex index bb24842..cbe34c4 100644 --- a/lib/ex_codecs/native.ex +++ b/lib/ex_codecs/native.ex @@ -22,6 +22,9 @@ defmodule ExCodecs.Native do otp_app: :ex_codecs, crate: :ex_codecs_native, version: version, + # This URL must match the GitHub repo that hosts release artifacts. + # If the repo is renamed, published checksums become invalid until + # a new release is cut. Keep in sync with @source_url in mix.exs. base_url: "https://github.com/thanos/codecs/releases/download/v#{version}", mode: :release, nif_versions: ["2.17"], @@ -136,6 +139,8 @@ defmodule ExCodecs.Native do def nif_loaded? do is_map(codec_versions()) rescue + # ErlangError: NIF not loaded (stub raises nif_error). + # ArgumentError: rare Rustler resource-arity mismatch on some OTPs. ErlangError -> false ArgumentError -> false end diff --git a/lib/ex_codecs/nif.ex b/lib/ex_codecs/nif.ex index ef60dc2..b41500d 100644 --- a/lib/ex_codecs/nif.ex +++ b/lib/ex_codecs/nif.ex @@ -1,13 +1,34 @@ defmodule ExCodecs.NIF do - @moduledoc false + @moduledoc """ + Shared NIF result wrapping and the library-wide decompress size ceiling. + + Compression codecs call `safe_call/2` around `ExCodecs.Native` functions so + callers always see `{:ok, binary()}` or `{:error, %ExCodecs.Error{}}`, never + raw `:erlang.nif_error/1` exceptions on the public API. + + ## Safety policy + + Decompression defaults to a **256 MiB** output ceiling (`default_max_output_size/0`). + Override per call with `max_output_size:` (bytes). Exceeding the limit returns + `:output_limit_exceeded`. + + ## Panic classification + + `safe_call/2` maps `:nif_not_loaded` precisely. Panic-like `ErlangError` + payloads are logged at warning and returned as `:compression_failed` with + message `"Native codec crashed"`. + """ - # Default decompress ceiling: 256 MiB. Override with `max_output_size:`. @default_max_output_size 268_435_456 - @doc false + @doc "Default decompress ceiling in bytes (256 MiB)." + @spec default_max_output_size() :: pos_integer() def default_max_output_size, do: @default_max_output_size - @doc false + @doc """ + Reads `:max_output_size` from `opts`, defaulting to `default_max_output_size/0`. + """ + @spec max_output_size(keyword()) :: {:ok, pos_integer()} | {:error, ExCodecs.Error.t()} def max_output_size(opts) when is_list(opts) do case Keyword.get(opts, :max_output_size, @default_max_output_size) do n when is_integer(n) and n > 0 -> @@ -22,10 +43,7 @@ defmodule ExCodecs.NIF do end @doc """ - Wraps raw NIF error tuples into ExCodecs.Error structs. - - NIF functions return `{:error, atom}` tuples. This helper converts - them into `{:error, %ExCodecs.Error{}}` tuples for consistent error handling. + Wraps raw NIF `{:ok, binary}` / `{:error, atom}` results into `ExCodecs.Error`. """ @spec wrap(atom(), {:ok, binary()} | {:error, atom()} | term()) :: {:ok, binary()} | {:error, ExCodecs.Error.t()} @@ -43,8 +61,7 @@ defmodule ExCodecs.NIF do ExCodecs.Error.new(:output_limit_exceeded, codec: codec, message: - "Decompressed output exceeded max_output_size " <> - "(default #{@default_max_output_size} bytes). " <> + "Decompressed output exceeded max_output_size. " <> "Pass a larger max_output_size: for trusted inputs." )} @@ -54,8 +71,23 @@ defmodule ExCodecs.NIF do def wrap(codec, {:error, :invalid_options}), do: {:error, ExCodecs.Error.new(:invalid_options, codec: codec)} - def wrap(codec, {:error, reason}) when is_atom(reason), - do: {:error, ExCodecs.Error.new(reason, codec: codec)} + def wrap(codec, {:error, reason}) when is_atom(reason) and reason not in [ + :compression_failed, + :decompression_failed, + :output_limit_exceeded, + :invalid_data, + :invalid_options, + :nif_not_loaded, + :io_error, + :truncated_input + ] do + {:error, + ExCodecs.Error.new(:invalid_data, + codec: codec, + message: "Unexpected NIF error atom: #{inspect(reason)}", + details: reason + )} + end def wrap(codec, other) do {:error, @@ -68,7 +100,7 @@ defmodule ExCodecs.NIF do @doc """ Invokes a zero-arity NIF function and wraps ErlangError/`nif_not_loaded`. """ - @spec safe_call(atom(), (-> term())) :: {:ok, term()} | {:error, ExCodecs.Error.t()} + @spec safe_call(atom(), (-> term())) :: {:ok, binary()} | {:error, ExCodecs.Error.t()} def safe_call(codec, fun) when is_function(fun, 0) do wrap(codec, fun.()) rescue @@ -78,10 +110,28 @@ defmodule ExCodecs.NIF do {:error, ExCodecs.Error.new(:nif_not_loaded, codec: codec)} other -> - {:error, ExCodecs.Error.new(:invalid_data, codec: codec, details: other)} + if nif_panic?(other) do + require Logger + + Logger.warning("ExCodecs NIF panic in #{inspect(codec)}: #{inspect(other, limit: 32)}") + + {:error, + ExCodecs.Error.new(:compression_failed, + codec: codec, + message: "Native codec crashed", + details: other + )} + else + {:error, ExCodecs.Error.new(:invalid_data, codec: codec, details: other)} + end end e in ArgumentError -> {:error, ExCodecs.Error.new(:invalid_data, codec: codec, message: Exception.message(e))} end + + defp nif_panic?(other) when is_binary(other), do: true + defp nif_panic?({:error, _}), do: true + defp nif_panic?(other) when is_tuple(other), do: true + defp nif_panic?(_), do: false end diff --git a/lib/ex_codecs/spatial.ex b/lib/ex_codecs/spatial.ex index 2ad2a6a..a7dd440 100644 --- a/lib/ex_codecs/spatial.ex +++ b/lib/ex_codecs/spatial.ex @@ -95,7 +95,7 @@ defmodule ExCodecs.Spatial do ## Arguments - * `format` (`atom()`) — the candidate format name. + * `format` (`atom()`) - the candidate format name. ## Returns @@ -127,10 +127,10 @@ defmodule ExCodecs.Spatial do ## Arguments - * `data` (`PointCloud.t() | GaussianCloud.t()`) — a point cloud for + * `data` (`PointCloud.t() | GaussianCloud.t()`) - a point cloud for `:ply` or `:spatial_binary`, or a Gaussian cloud for `:ply` or `:gsplat`. - * `opts` (`keyword()`) — options passed to the selected codec: - * `:format` — spatial container: `:ply` (default), `:spatial_binary`, or + * `opts` (`keyword()`) - options passed to the selected codec: + * `:format` - spatial container: `:ply` (default), `:spatial_binary`, or `:gsplat`. This key is removed before codec dispatch. * for PLY, `:ply_format` selects `:ascii`, `:binary`, `:binary_le`, or `:binary_be`; `:comments` overrides metadata comments. @@ -214,10 +214,10 @@ defmodule ExCodecs.Spatial do ## Arguments - * `data` (`binary()`) — a complete encoded payload. - * `opts` (`keyword()`) — decode options: - * `:format` — `:ply` (default), `:spatial_binary`, or `:gsplat`. - * `:as` — PLY interpretation: `:auto` (default), `:point_cloud`, or + * `data` (`binary()`) - a complete encoded payload. + * `opts` (`keyword()`) - decode options: + * `:format` - `:ply` (default), `:spatial_binary`, or `:gsplat`. + * `:as` - PLY interpretation: `:auto` (default), `:point_cloud`, or `:gaussian_cloud`. In `:auto`, Gaussian property names select a `%GaussianCloud{}`; other vertex schemas select a `%PointCloud{}`. @@ -287,10 +287,10 @@ defmodule ExCodecs.Spatial do ## Arguments - * `source` (`Path.t() | binary()`) — a filesystem path or encoded payload. + * `source` (`Path.t() | binary()`) - a filesystem path or encoded payload. Since paths are binaries in Elixir, use `source: :file` to force path interpretation or `source: :binary` to force payload interpretation. - * `opts` (`keyword()`) — requires `:format` (`:ply`, + * `opts` (`keyword()`) - requires `:format` (`:ply`, `:spatial_binary`, or `:gsplat`). `:source` may be `:auto` (default), `:file`, or `:binary`; remaining options go to the selected decoder. @@ -335,8 +335,8 @@ defmodule ExCodecs.Spatial do ## Arguments - * `enumerable` (`Enumerable.t()`) — `%Point{}` or `%Gaussian{}` elements. - * `opts` (`keyword()`) — requires `:format` (`:ply`, + * `enumerable` (`Enumerable.t()`) - `%Point{}` or `%Gaussian{}` elements. + * `opts` (`keyword()`) - requires `:format` (`:ply`, `:spatial_binary`, or `:gsplat`); remaining options are the same as `encode/2`. diff --git a/lib/ex_codecs/spatial/accel.ex b/lib/ex_codecs/spatial/accel.ex index 3e38fb0..73c6320 100644 --- a/lib/ex_codecs/spatial/accel.ex +++ b/lib/ex_codecs/spatial/accel.ex @@ -1,13 +1,49 @@ defmodule ExCodecs.Spatial.Accel do - @moduledoc false + @moduledoc """ + Elixir facade over the spatial Rust NIFs (EXCP, GSPL, binary PLY). - # Thin facade over DirtyCpu spatial NIFs. Callers fall back to pure Elixir - # when `available?/0` is false or a call returns `:nif_not_loaded`. + Callers fall back to pure Elixir when `available?/0` is false or a call + returns `:nif_not_loaded`. This module defines the **cross-language ABI**: + row tuple shapes must match `native/ex_codecs_native/src/spatial.rs`. + + ## Chunk semantics + + Unpack NIFs return `{:ok, {rows, next_offset}}`. Pass at most `chunk_size/0` + (4096) rows per call and advance `offset` from `next_offset`. Whole-payload + decode paths loop the same way. A partial final record is silently skipped + (the unpack `break`s); callers can detect truncation by comparing + `next_offset` to the expected body size. + + ## Data coercion + + Out-of-range color bytes are **clamped** to `[0, 255]` (not wrapped). Missing + SH coefficients are zero-filled to the declared `sh_rest` length. Both + coercions are lossy: a round-trip through encode may not reproduce the + original if the input contained out-of-range values. + + ## Row formats + + * EXCP point row: `{x, y, z, color, normal}` where `color` / `normal` are + `nil` or tuples matching header flags + * GSPL gaussian row: + `{{x, y, z}, {r, g, b}, opacity, {sx, sy, sz}, {rw, rx, ry, rz}, sh_list}` + * PLY binary: flat numeric lists per vertex, typed by `ply_type_tag/1` + + ## mmap + + Do not truncate or replace a file while an mmap resource from `mmap_open/1` + is live — concurrent truncation can SIGBUS and kill the BEAM. Prefer + `accel: false` / plain IO for untrusted paths. + + This module is intentional infrastructure; HexDocs may still group it under + internals depending on ExDoc filters. + """ alias ExCodecs.Native alias ExCodecs.Spatial.{Gaussian, Point} @default_chunk 4_096 + @available_key {__MODULE__, :available} @ply_type_tags %{ char: 1, @@ -20,7 +56,23 @@ defmodule ExCodecs.Spatial.Accel do double: 8 } + @doc """ + Returns whether spatial NIFs are loaded (cached for the VM lifetime). + """ + @spec available?() :: boolean() def available? do + case :persistent_term.get(@available_key, :unset) do + :unset -> + value = probe_available() + :persistent_term.put(@available_key, value) + value + + value when is_boolean(value) -> + value + end + end + + defp probe_available do case safe(fn -> Native.codec_versions() end) do map when is_map(map) -> Map.has_key?(map, "spatial") or Map.has_key?(map, :spatial) @@ -32,6 +84,8 @@ defmodule ExCodecs.Spatial.Accel do end end + @doc "Preferred max row count per unpack NIF call." + @spec chunk_size() :: pos_integer() def chunk_size, do: @default_chunk # --- EXCP ----------------------------------------------------------------- @@ -86,7 +140,7 @@ defmodule ExCodecs.Spatial.Accel do def ply_binary_unpack(data, types, endian, offset \\ 0, max_count \\ @default_chunk) when is_binary(data) and is_list(types) do tags = Enum.map(types, &ply_type_tag/1) - little? = endian in [:binary_le, :little, true] + little? = endian in [:binary_le, :little] nif_chunk( fn -> Native.ply_binary_unpack(data, tags, little?, offset, max_count) end, @@ -96,7 +150,7 @@ defmodule ExCodecs.Spatial.Accel do def ply_binary_unpack_mmap(resource, types, endian, offset \\ 0, max_count \\ @default_chunk) do tags = Enum.map(types, &ply_type_tag/1) - little? = endian in [:binary_le, :little, true] + little? = endian in [:binary_le, :little] nif_chunk( fn -> Native.ply_binary_unpack_mmap(resource, tags, little?, offset, max_count) end, @@ -154,8 +208,9 @@ defmodule ExCodecs.Spatial.Accel do sh = case g.sh do nil -> [] - [_dc | rest] -> List.flatten(rest) - list when is_list(list) -> list |> List.flatten() |> Enum.drop(3) + [] -> [] + [[_ | _] | rest] -> List.flatten(rest) + [h | _] = list when is_number(h) -> list |> List.flatten() |> Enum.drop(3) end |> then(fn vals -> vals @@ -225,18 +280,21 @@ defmodule ExCodecs.Spatial.Accel do fun.() rescue e -> - # Rustler may raise ArgumentError for bad resources; that is not always an - # `%ErlangError{}`, so discriminate on the struct rather than `rescue in`. case e do - # coveralls-ignore-start %ErlangError{original: :nif_not_loaded} -> {:error, :nif_not_loaded} %ErlangError{original: other} -> {:error, other} - # coveralls-ignore-stop + %ArgumentError{} -> + {:error, Exception.message(e)} + other -> + require Logger + + Logger.warning("ExCodecs.Spatial.Accel unexpected exception: #{inspect(other)}") + {:error, Exception.message(other)} end end diff --git a/lib/ex_codecs/spatial/bounds.ex b/lib/ex_codecs/spatial/bounds.ex index 13217af..11fb2ca 100644 --- a/lib/ex_codecs/spatial/bounds.ex +++ b/lib/ex_codecs/spatial/bounds.ex @@ -65,7 +65,9 @@ defmodule ExCodecs.Spatial.Bounds do {number(), number(), number()}, {number(), number(), number()} ) :: t() - def new({min_x, min_y, min_z}, {max_x, max_y, max_z}) do + def new({min_x, min_y, min_z}, {max_x, max_y, max_z}) + when is_number(min_x) and is_number(min_y) and is_number(min_z) and + is_number(max_x) and is_number(max_y) and is_number(max_z) do %__MODULE__{ min_x: min_x * 1.0, min_y: min_y * 1.0, diff --git a/lib/ex_codecs/spatial/codec/binary.ex b/lib/ex_codecs/spatial/codec/binary.ex index b5ee79d..516dd33 100644 --- a/lib/ex_codecs/spatial/codec/binary.ex +++ b/lib/ex_codecs/spatial/codec/binary.ex @@ -13,11 +13,12 @@ defmodule ExCodecs.Spatial.Codec.Binary do Each record always has `x,y,z` as `f32`. Optional `rgb`/`rgba` as `u8` and normals as `f32` follow according to flags. - Generic attributes are not stored in this format — use PLY for that. + Generic attributes are not stored in this format - use PLY for that. """ alias ExCodecs.Error alias ExCodecs.Spatial.Accel + alias ExCodecs.Spatial.Codec.StreamSource alias ExCodecs.Spatial.{Metadata, Point, PointCloud} @magic "EXCP" @@ -42,10 +43,10 @@ defmodule ExCodecs.Spatial.Codec.Binary do ## Arguments - * `data` (`PointCloud.t()`) — a cloud whose `points` are valid `%Point{}` + * `data` (`PointCloud.t()`) - a cloud whose `points` are valid `%Point{}` structs with numeric coordinates, optional `{r, g, b}` or `{r, g, b, a}` color, and optional `{nx, ny, nz}` normal. - * `opts` (`keyword()`) — reserved and currently ignored. + * `opts` (`keyword()`) - reserved and currently ignored. ## Returns @@ -74,19 +75,22 @@ defmodule ExCodecs.Spatial.Codec.Binary do def encode(data, opts \\ []) def encode(%PointCloud{points: points}, opts) do - has_color? = Enum.any?(points, &Point.colored?/1) - has_alpha? = Enum.any?(points, fn p -> match?({_, _, _, _}, p.color) end) - has_normal? = Enum.any?(points, &Point.has_normal?/1) - - flags = - 0 - |> then(fn f -> if has_color?, do: Bitwise.bor(f, @flag_color), else: f end) - |> then(fn f -> if has_alpha?, do: Bitwise.bor(f, @flag_alpha), else: f end) - |> then(fn f -> if has_normal?, do: Bitwise.bor(f, @flag_normal), else: f end) + {flags, count} = + Enum.reduce(points, {0, 0}, fn p, {f, n} -> + f = + case p.color do + {_, _, _, _} -> Bitwise.bor(f, Bitwise.bor(@flag_color, @flag_alpha)) + {_, _, _} -> Bitwise.bor(f, @flag_color) + nil -> f + end + + f = if Point.has_normal?(p), do: Bitwise.bor(f, @flag_normal), else: f + {f, n + 1} + end) header = <<@magic::binary, @version::little-unsigned-16, flags::little-unsigned-16, - length(points)::little-unsigned-64>> + count::little-unsigned-64>> body = encode_points_body(points, flags, opts) {:ok, header <> body} @@ -108,9 +112,9 @@ defmodule ExCodecs.Spatial.Codec.Binary do ## Arguments - * `enumerable` (`Enumerable.t()`) — `%Point{}` elements. - * `path` (`Path.t()`) — destination file path. - * `opts` (`keyword()`) — requires `:schema`, a list such as `[]`, + * `enumerable` (`Enumerable.t()`) - `%Point{}` elements. + * `path` (`Path.t()`) - destination file path. + * `opts` (`keyword()`) - requires `:schema`, a list such as `[]`, `[:color]`, `[:color, :alpha]`, `[:normal]`, or `[:color, :normal]`. Map form `%{color: true, alpha: false, normal: true}` is also accepted. `:alpha` implies color bytes (RGBA). @@ -145,11 +149,15 @@ defmodule ExCodecs.Spatial.Codec.Binary do with {:ok, flags} <- fetch_schema_flags(opts), {:ok, io} <- open_write(path) do try do - :ok = IO.binwrite(io, excp_header(flags, 0)) - count = write_point_chunks(enumerable, io, flags, opts) - {:ok, 0} = :file.position(io, 0) - :ok = IO.binwrite(io, excp_header(flags, count)) - :ok + with :ok <- binwrite(io, excp_header(flags, 0)), + {:ok, count} <- write_point_chunks(enumerable, io, flags, opts), + {:ok, _} <- :file.position(io, 0), + :ok <- binwrite(io, excp_header(flags, count)) do + :ok + else + {:error, %Error{}} = err -> err + {:error, reason} -> io_error(reason) + end catch {:bad_point, other} -> {:error, @@ -247,8 +255,8 @@ defmodule ExCodecs.Spatial.Codec.Binary do ## Arguments - * `data` (`binary()`) — a complete EXCP payload. - * `opts` (`keyword()`) — reserved and currently ignored. + * `data` (`binary()`) - a complete EXCP payload. + * `opts` (`keyword()`) - reserved and currently ignored. ## Returns @@ -287,10 +295,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do message: "Unsupported binary point format version #{version}" )} else - with {:ok, points, _} <- decode_points(rest, count, flags, opts) do - meta = Metadata.new(entries: %{"format" => "excp", "version" => version}) - {:ok, PointCloud.new(points, metadata: meta)} - end + decode_v1(flags, count, rest, version, opts) end end @@ -302,25 +307,44 @@ defmodule ExCodecs.Spatial.Codec.Binary do )} end + defp decode_v1(flags, count, rest, version, opts) do + known_flags = Bitwise.bor(@flag_color, Bitwise.bor(@flag_alpha, @flag_normal)) + + if Bitwise.band(flags, Bitwise.bnot(known_flags)) != 0 do + {:error, + Error.new(:invalid_data, + codec: :spatial_binary, + message: "Unknown EXCP flag bits: 0x#{Integer.to_string(flags, 16)}" + )} + else + with {:ok, points, _} <- decode_points(rest, count, flags, opts) do + meta = Metadata.new(entries: %{"format" => "excp", "version" => version}) + {:ok, PointCloud.new(points, metadata: meta)} + end + end + end + @doc """ Returns an enumerable over points in an EXCP payload or file. ## Streaming behavior * `source: :file` (or `:auto` path detection): header then chunked - records — Rust mmap + DirtyCpu unpack when available, otherwise + records - Rust mmap + DirtyCpu unpack when available, otherwise `IO.binread/2` per record. * `source: :binary`: chunked Rust unpack when available; otherwise materializes through `decode/2`. - Prefer `source: :file` for multi‑MB / multi‑GB `.excp` paths. + Prefer `source: :file` for multi-MB / multi-GB `.excp` paths. Pass `accel: false` to force the pure-Elixir path. ## Arguments - * `source` (`Path.t() | binary()`) — filesystem path or complete EXCP + * `source` (`Path.t() | binary()`) - filesystem path or complete EXCP payload. Both are binaries, so `:source` controls resolution. - * `opts` (`keyword()`) — `:source` may be `:auto` (default), `:file`, or + * `opts` (`keyword()`) - `:source` may be `:binary` (default), `:file`, or + `:auto`. Prefer `:file` for paths and `:binary` for payloads; `:auto` is + opt-in path sniffing (see `docs/spatial_formats.md`). `:binary`. ## Returns @@ -328,8 +352,8 @@ defmodule ExCodecs.Spatial.Codec.Binary do An `Enumerable.t()` yielding `%Point{}` values. Failures are delayed until enumeration as exactly one `{:error, %ExCodecs.Error{}}`: - * `reason: :io_error` — the file cannot be opened or read - * `reason: :invalid_data` — bad magic/version, truncated header, or a + * `reason: :io_error` - the file cannot be opened or read + * `reason: :invalid_data` - bad magic/version, truncated header, or a truncated record mid-stream (`codec: :spatial_binary`) ## Raises / exceptions @@ -348,7 +372,7 @@ defmodule ExCodecs.Spatial.Codec.Binary do def stream_decode(source, opts \\ []) def stream_decode(data, opts) when is_binary(data) do - case resolve_source(data, opts) do + case StreamSource.resolve(data, opts, :spatial_binary, &path_like?/1) do {:ok, :binary, bin} -> stream_from_binary(bin, opts) @@ -358,6 +382,9 @@ defmodule ExCodecs.Spatial.Codec.Binary do &next_excp_item/1, &close_excp_file/1 ) + + {:error, error} -> + error_stream(error) end end @@ -401,27 +428,8 @@ defmodule ExCodecs.Spatial.Codec.Binary do ) end - defp resolve_source(bin, opts) do - case Keyword.get(opts, :source, :auto) do - :binary -> - {:ok, :binary, bin} - - :file -> - {:ok, :file, bin} - - :auto -> - if path_like?(bin) and File.regular?(bin) do - {:ok, :file, bin} - else - {:ok, :binary, bin} - end - end - end - defp path_like?(bin) do - byte_size(bin) < 4096 and not String.starts_with?(bin, @magic) and - (String.contains?(bin, "/") or String.contains?(bin, "\\") or - String.ends_with?(bin, [".excp", ".bin"])) + StreamSource.path_like?(bin, &String.starts_with?(&1, @magic), [".excp", ".bin"]) end defp open_excp_file(path, opts) do @@ -765,19 +773,21 @@ defmodule ExCodecs.Spatial.Codec.Binary do defp normalize_rgb(color) do case color do nil -> {0, 0, 0} - {r, g, b} -> {trunc(r), trunc(g), trunc(b)} - {r, g, b, _a} -> {trunc(r), trunc(g), trunc(b)} + {r, g, b} -> {clamp_byte(r), clamp_byte(g), clamp_byte(b)} + {r, g, b, _a} -> {clamp_byte(r), clamp_byte(g), clamp_byte(b)} end end defp normalize_rgba(color) do case color do nil -> {0, 0, 0, 255} - {r, g, b} -> {trunc(r), trunc(g), trunc(b), 255} - {r, g, b, a} -> {trunc(r), trunc(g), trunc(b), trunc(a)} + {r, g, b} -> {clamp_byte(r), clamp_byte(g), clamp_byte(b), 255} + {r, g, b, a} -> {clamp_byte(r), clamp_byte(g), clamp_byte(b), clamp_byte(a)} end end + defp clamp_byte(v), do: min(max(trunc(v), 0), 255) + defp encode_points_body(points, flags, opts) do if accel?(opts) do case Accel.excp_pack(points, flags) do @@ -803,10 +813,13 @@ defmodule ExCodecs.Spatial.Codec.Binary do enumerable |> Stream.chunk_every(chunk_size) - |> Enum.reduce(0, fn chunk, n -> + |> Enum.reduce_while({:ok, 0}, fn chunk, {:ok, n} -> points = assert_points!(chunk) - :ok = write_points_chunk(io, points, flags, opts) - n + length(points) + + case write_points_chunk(io, points, flags, opts) do + :ok -> {:cont, {:ok, n + length(points)}} + {:error, _} = err -> {:halt, err} + end end) end @@ -821,40 +834,69 @@ defmodule ExCodecs.Spatial.Codec.Binary do if accel?(opts) do case Accel.excp_pack(points, flags) do {:ok, bin} -> - IO.binwrite(io, bin) + binwrite(io, bin) # coveralls-ignore-start _ -> - Enum.each(points, &IO.binwrite(io, encode_point(&1, flags))) + write_points_elixir(io, points, flags) # coveralls-ignore-stop end else - Enum.each(points, &IO.binwrite(io, encode_point(&1, flags))) + write_points_elixir(io, points, flags) end + end + + defp write_points_elixir(io, points, flags) do + Enum.reduce_while(points, :ok, fn p, :ok -> + case binwrite(io, encode_point(p, flags)) do + :ok -> {:cont, :ok} + {:error, _} = err -> {:halt, err} + end + end) + end - :ok + defp binwrite(io, data) do + case :file.write(io, data) do + :ok -> :ok + {:error, reason} -> io_error(reason) + end end - defp accel?(opts), do: Keyword.get(opts, :accel, true) != false and Accel.available?() + defp accel?(opts), do: StreamSource.accel?(opts) defp decode_points(bin, 0, _flags, _opts), do: {:ok, [], bin} defp decode_points(bin, count, flags, opts) do if accel?(opts) do - case Accel.excp_unpack(bin, flags, 0, count) do - {:ok, {points, _}} when length(points) == count -> - {:ok, points, <<>>} - - # Incomplete / failed Accel unpack: Elixir path preserves field-specific - # truncation messages (RGB / RGBA / normal). - _ -> - decode_points_elixir(bin, count, flags) - end + decode_points_accel(bin, count, flags, 0, 0, []) else decode_points_elixir(bin, count, flags) end end + defp decode_points_accel(_bin, count, _flags, _offset, i, acc) when i >= count do + {:ok, Enum.reverse(acc), <<>>} + end + + defp decode_points_accel(bin, count, flags, offset, i, acc) do + want = min(Accel.chunk_size(), count - i) + + case Accel.excp_unpack(bin, flags, offset, want) do + {:ok, {points, next}} when length(points) == want -> + decode_points_accel(bin, count, flags, next, i + want, Enum.reverse(points, acc)) + + # Incomplete / failed Accel unpack: Elixir path preserves field-specific + # truncation messages (RGB / RGBA / normal), resuming at remaining bytes. + _ -> + rest = binary_part(bin, offset, byte_size(bin) - offset) + + case decode_points_elixir(rest, count - i, flags) do + {:ok, points, leftover} -> {:ok, Enum.reverse(acc, points), leftover} + {:error, _} = err -> err + end + end + end + defp decode_points_elixir(bin, count, flags) do Enum.reduce_while(1..count, {:ok, [], bin}, fn _, {:ok, acc, rest} -> case decode_point(rest, flags) do diff --git a/lib/ex_codecs/spatial/codec/gsplat.ex b/lib/ex_codecs/spatial/codec/gsplat.ex index f104383..371e3de 100644 --- a/lib/ex_codecs/spatial/codec/gsplat.ex +++ b/lib/ex_codecs/spatial/codec/gsplat.ex @@ -23,6 +23,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do alias ExCodecs.Error alias ExCodecs.Spatial.Accel + alias ExCodecs.Spatial.Codec.StreamSource alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Metadata} @magic "GSPL" @@ -45,11 +46,11 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do ## Arguments - * `data` (`GaussianCloud.t()`) — a cloud of valid `%Gaussian{}` structs. + * `data` (`GaussianCloud.t()`) - a cloud of valid `%Gaussian{}` structs. Each Gaussian has position/color/scale 3-tuples, a rotation 4-tuple, numeric opacity, and optional SH data shaped as `[dc | rest]` (nested rest lists are flattened). - * `opts` (`keyword()`) — reserved and currently ignored. + * `opts` (`keyword()`) - reserved and currently ignored. ## Returns @@ -102,9 +103,9 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do ## Arguments - * `enumerable` (`Enumerable.t()`) — `%Gaussian{}` elements. - * `path` (`Path.t()`) — destination file path. - * `opts` (`keyword()`) — requires `:schema`. Use `[]` / `%{}` for no SH + * `enumerable` (`Enumerable.t()`) - `%Gaussian{}` elements. + * `path` (`Path.t()`) - destination file path. + * `opts` (`keyword()`) - requires `:schema`. Use `[]` / `%{}` for no SH rest coefficients, or `[sh_rest: n]` / `%{sh_rest: n}` for `n` shared rest floats per Gaussian (shorter lists are zero-padded). @@ -140,13 +141,15 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do flags = if sh_rest > 0, do: 1, else: 0 try do - :ok = IO.binwrite(io, gspl_header(flags, 0, sh_rest)) - - count = write_gaussian_chunks(enumerable, io, sh_rest, opts) - - {:ok, 0} = :file.position(io, 0) - :ok = IO.binwrite(io, gspl_header(flags, count, sh_rest)) - :ok + with :ok <- binwrite(io, gspl_header(flags, 0, sh_rest)), + {:ok, count} <- write_gaussian_chunks(enumerable, io, sh_rest, opts), + {:ok, _} <- :file.position(io, 0), + :ok <- binwrite(io, gspl_header(flags, count, sh_rest)) do + :ok + else + {:error, %Error{}} = err -> err + {:error, reason} -> io_error(reason) + end catch {:bad_gaussian, other} -> {:error, @@ -237,8 +240,8 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do ## Arguments - * `data` (`binary()`) — a complete GSPL payload. - * `opts` (`keyword()`) — reserved and currently ignored. + * `data` (`binary()`) - a complete GSPL payload. + * `opts` (`keyword()`) - reserved and currently ignored. ## Returns @@ -294,19 +297,21 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do ## Streaming behavior * `source: :file` (or `:auto` path detection): header then chunked - records — Rust mmap + DirtyCpu unpack when available, otherwise + records - Rust mmap + DirtyCpu unpack when available, otherwise `IO.binread/2` per record. * `source: :binary`: chunked Rust unpack when available; otherwise materializes through `decode/2`. - Prefer `source: :file` for multi‑MB / multi‑GB `.gspl` paths. + Prefer `source: :file` for multi-MB / multi-GB `.gspl` paths. Pass `accel: false` to force the pure-Elixir path. ## Arguments - * `source` (`Path.t() | binary()`) — filesystem path or complete GSPL + * `source` (`Path.t() | binary()`) - filesystem path or complete GSPL payload. Both are binaries, so `:source` controls resolution. - * `opts` (`keyword()`) — `:source` may be `:auto` (default), `:file`, or + * `opts` (`keyword()`) - `:source` may be `:binary` (default), `:file`, or + `:auto`. Prefer `:file` for paths and `:binary` for payloads; `:auto` is + opt-in path sniffing (see `docs/spatial_formats.md`). `:binary`. ## Returns @@ -314,8 +319,8 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do An `Enumerable.t()` yielding `%Gaussian{}` values. Failures are delayed until enumeration as exactly one `{:error, %ExCodecs.Error{}}`: - * `reason: :io_error` — the file cannot be opened or read - * `reason: :invalid_data` — bad magic/version, truncated header, or a + * `reason: :io_error` - the file cannot be opened or read + * `reason: :invalid_data` - bad magic/version, truncated header, or a truncated record mid-stream (`codec: :gsplat`) ## Raises / exceptions @@ -334,7 +339,7 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do def stream_decode(source, opts \\ []) def stream_decode(data, opts) when is_binary(data) do - case resolve_source(data, opts) do + case StreamSource.resolve(data, opts, :gsplat, &path_like?/1) do {:ok, :binary, bin} -> stream_from_binary(bin, opts) @@ -344,6 +349,9 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do &next_gspl_item/1, &close_gspl_file/1 ) + + {:error, error} -> + error_stream(error) end end @@ -385,27 +393,8 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do error_stream(Error.new(:invalid_data, codec: :gsplat, message: "Invalid GSPLAT binary")) end - defp resolve_source(bin, opts) do - case Keyword.get(opts, :source, :auto) do - :binary -> - {:ok, :binary, bin} - - :file -> - {:ok, :file, bin} - - :auto -> - if path_like?(bin) and File.regular?(bin) do - {:ok, :file, bin} - else - {:ok, :binary, bin} - end - end - end - defp path_like?(bin) do - byte_size(bin) < 4096 and not String.starts_with?(bin, @magic) and - (String.contains?(bin, "/") or String.contains?(bin, "\\") or - String.ends_with?(bin, [".gspl", ".bin"])) + StreamSource.path_like?(bin, &String.starts_with?(&1, @magic), [".gspl", ".bin"]) end defp open_gspl_file(path, opts) do @@ -688,8 +677,9 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do gaussians |> Enum.map(fn %{sh: nil} -> 0 - %{sh: [_dc | rest]} -> length(List.flatten(rest)) - %{sh: other} when is_list(other) -> max(length(List.flatten(other)) - 3, 0) + %{sh: []} -> 0 + %{sh: [[_ | _] | rest]} -> length(List.flatten(rest)) + %{sh: [h | _] = other} when is_number(h) -> max(length(List.flatten(other)) - 3, 0) end) |> Enum.max(fn -> 0 end) end @@ -721,10 +711,9 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do end defp sh_rest_values(nil), do: [] - defp sh_rest_values([_dc | rest]), do: List.flatten(rest) - # coveralls-ignore-start - defp sh_rest_values(list) when is_list(list), do: list |> List.flatten() |> Enum.drop(3) - # coveralls-ignore-stop + defp sh_rest_values([]), do: [] + defp sh_rest_values([[_ | _] | rest]), do: List.flatten(rest) + defp sh_rest_values([h | _] = list) when is_number(h), do: list |> List.flatten() |> Enum.drop(3) defp encode_gaussians_body(gaussians, sh_rest, opts) do if accel?(opts) do @@ -751,10 +740,13 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do enumerable |> Stream.chunk_every(chunk_size) - |> Enum.reduce(0, fn chunk, n -> + |> Enum.reduce_while({:ok, 0}, fn chunk, {:ok, n} -> gaussians = assert_gaussians!(chunk) - :ok = write_gaussian_chunk(io, gaussians, sh_rest, opts) - n + length(gaussians) + + case write_gaussian_chunk(io, gaussians, sh_rest, opts) do + :ok -> {:cont, {:ok, n + length(gaussians)}} + {:error, _} = err -> {:halt, err} + end end) end @@ -769,40 +761,76 @@ defmodule ExCodecs.Spatial.Codec.Gsplat do if accel?(opts) do case Accel.gspl_pack(gaussians, sh_rest) do {:ok, bin} -> - IO.binwrite(io, bin) + binwrite(io, bin) # coveralls-ignore-start _ -> - Enum.each(gaussians, &IO.binwrite(io, encode_gaussian(&1, sh_rest))) + write_gaussians_elixir(io, gaussians, sh_rest) # coveralls-ignore-stop end else - Enum.each(gaussians, &IO.binwrite(io, encode_gaussian(&1, sh_rest))) + write_gaussians_elixir(io, gaussians, sh_rest) end + end - :ok + defp write_gaussians_elixir(io, gaussians, sh_rest) do + Enum.reduce_while(gaussians, :ok, fn g, :ok -> + case binwrite(io, encode_gaussian(g, sh_rest)) do + :ok -> {:cont, :ok} + {:error, _} = err -> {:halt, err} + end + end) + end + + defp binwrite(io, data) do + case :file.write(io, data) do + :ok -> :ok + {:error, reason} -> io_error(reason) + end end - defp accel?(opts), do: Keyword.get(opts, :accel, true) != false and Accel.available?() + defp accel?(opts), do: StreamSource.accel?(opts) defp decode_gaussians(bin, 0, _sh_rest, _opts), do: {:ok, [], bin} defp decode_gaussians(bin, count, sh_rest, opts) do if accel?(opts) do - case Accel.gspl_unpack(bin, sh_rest, 0, count) do - {:ok, {gaussians, _}} when length(gaussians) == count -> - {:ok, gaussians, <<>>} - - # Incomplete / failed Accel unpack: Elixir path preserves "Truncated SH" - # and other field-specific messages. - _ -> - decode_gaussians_elixir(bin, count, sh_rest) - end + decode_gaussians_accel(bin, count, sh_rest, 0, 0, []) else decode_gaussians_elixir(bin, count, sh_rest) end end + defp decode_gaussians_accel(_bin, count, _sh_rest, _offset, i, acc) when i >= count do + {:ok, Enum.reverse(acc), <<>>} + end + + defp decode_gaussians_accel(bin, count, sh_rest, offset, i, acc) do + want = min(Accel.chunk_size(), count - i) + + case Accel.gspl_unpack(bin, sh_rest, offset, want) do + {:ok, {gaussians, next}} when length(gaussians) == want -> + decode_gaussians_accel( + bin, + count, + sh_rest, + next, + i + want, + Enum.reverse(gaussians, acc) + ) + + # Incomplete / failed Accel unpack: Elixir path preserves "Truncated SH" + # and other field-specific messages, resuming at remaining bytes. + _ -> + rest = binary_part(bin, offset, byte_size(bin) - offset) + + case decode_gaussians_elixir(rest, count - i, sh_rest) do + {:ok, gs, leftover} -> {:ok, Enum.reverse(acc, gs), leftover} + {:error, _} = err -> err + end + end + end + defp decode_gaussians_elixir(bin, count, sh_rest) do Enum.reduce_while(1..count, {:ok, [], bin}, fn _, {:ok, acc, rest} -> case decode_gaussian(rest, sh_rest) do diff --git a/lib/ex_codecs/spatial/codec/ply.ex b/lib/ex_codecs/spatial/codec/ply.ex index eebac3f..7ed510d 100644 --- a/lib/ex_codecs/spatial/codec/ply.ex +++ b/lib/ex_codecs/spatial/codec/ply.ex @@ -4,17 +4,18 @@ defmodule ExCodecs.Spatial.Codec.PLY do ## Options - * `:format` / `:ply_format` — wire encoding: `:ascii` (default), + * `:format` / `:ply_format` - wire encoding: `:ascii` (default), `:binary` / `:binary_le`, `:binary_be` (not Spatial's `format: :ply`) - * `:comments` — header comments - * `:as` — decode as `:auto` | `:point_cloud` | `:gaussian_cloud` - * `:source` — for `stream_decode/2`: `:auto` | `:file` | `:binary` + * `:comments` - header comments + * `:as` - decode as `:auto` | `:point_cloud` | `:gaussian_cloud` + * `:source` - for `stream_decode/2`: `:auto` | `:file` | `:binary` Public functions: `encode/2`, `decode/2`, `stream_decode/2`. """ alias ExCodecs.Error alias ExCodecs.Spatial.Accel + alias ExCodecs.Spatial.Codec.StreamSource alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Metadata, Point, PointCloud} @typedoc """ @@ -46,22 +47,24 @@ defmodule ExCodecs.Spatial.Codec.PLY do ## Arguments - * `data` (`PointCloud.t() | GaussianCloud.t()`) — a cloud containing valid + * `data` (`PointCloud.t() | GaussianCloud.t()`) - a cloud containing valid `%Point{}` or `%Gaussian{}` structs. - * `opts` (`keyword()`) — supported options: - * `:ply_format` — preferred wire format: `:ascii` (default), `:binary` + * `opts` (`keyword()`) - supported options: + * `:ply_format` - preferred wire format: `:ascii` (default), `:binary` (little-endian), `:binary_le`, or `:binary_be`. - * `:format` — legacy alias for `:ply_format` when calling this codec + * `:format` - legacy alias for `:ply_format` when calling this codec directly. `:ply_format` takes precedence. - * `:comments` — enumerable of strings for PLY `comment` header lines; + * `:comments` - enumerable of strings for PLY `comment` header lines; defaults to `cloud.metadata.comments`. `:little` and `:big` are also normalized to little-/big-endian output. - Any unrecognized wire-format value currently falls back to `:ascii`. + Unrecognized wire-format values return `:invalid_options`. ## Returns * `{:ok, payload}` where `payload` is an ASCII or binary PLY `binary()`. + * `{:error, %ExCodecs.Error{reason: :invalid_options, codec: :ply}}` when + `:format` / `:ply_format` is not a supported wire encoding. * `{:error, %ExCodecs.Error{reason: :invalid_data, codec: :ply}}` when `data` is not a supported cloud or any exception occurs while deriving properties, building the header, or encoding rows. @@ -87,16 +90,17 @@ defmodule ExCodecs.Spatial.Codec.PLY do def encode(data, opts \\ []) def encode(%PointCloud{} = cloud, opts) do - format = normalize_format(ply_format_opt(opts)) comments = Keyword.get(opts, :comments, cloud.metadata.comments) - with {:ok, props} <- point_properties(cloud.points) do + with {:ok, format} <- normalize_format(ply_format_opt(opts)), + :ok <- validate_comments(comments), + {:ok, props} <- point_properties(cloud.points) do body = encode_points(cloud.points, props, format) header = build_header(format, length(cloud.points), props, comments) {:ok, header <> body} end rescue - e -> + e in [ArithmeticError, FunctionClauseError, ArgumentError] -> {:error, Error.new(:invalid_data, codec: :ply, @@ -105,14 +109,17 @@ defmodule ExCodecs.Spatial.Codec.PLY do end def encode(%GaussianCloud{} = cloud, opts) do - format = normalize_format(ply_format_opt(opts)) comments = Keyword.get(opts, :comments, cloud.metadata.comments) - props = gaussian_properties(cloud.gaussians) - body = encode_gaussians(cloud.gaussians, props, format) - header = build_header(format, length(cloud.gaussians), props, comments) - {:ok, header <> body} + + with {:ok, format} <- normalize_format(ply_format_opt(opts)), + :ok <- validate_comments(comments) do + props = gaussian_properties(cloud.gaussians) + body = encode_gaussians(cloud.gaussians, props, format) + header = build_header(format, length(cloud.gaussians), props, comments) + {:ok, header <> body} + end rescue - e -> + e in [ArithmeticError, FunctionClauseError, ArgumentError] -> {:error, Error.new(:invalid_data, codec: :ply, @@ -144,8 +151,8 @@ defmodule ExCodecs.Spatial.Codec.PLY do ## Arguments - * `data` (`binary()`) — the complete PLY header and vertex body. - * `opts` (`keyword()`) — `:as` may be `:auto` (default), `:point_cloud`, or + * `data` (`binary()`) - the complete PLY header and vertex body. + * `opts` (`keyword()`) - `:as` may be `:auto` (default), `:point_cloud`, or `:gaussian_cloud`. `:auto` chooses Gaussian output when the property set contains `f_dc_0`, `scale_0`, or `rot_0`; otherwise it chooses points. The wire format is always read from the header. @@ -156,16 +163,16 @@ defmodule ExCodecs.Spatial.Codec.PLY do * `{:error, %ExCodecs.Error{reason: :invalid_data, codec: :ply}}` for non-binary input, missing/invalid magic or format/header lines, missing vertex elements, invalid counts/properties, unsupported list/property - types, too few ASCII rows, or a short binary body. + types, invalid ASCII scalar tokens, token/property count mismatches, + missing `x`/`y`/`z`, too few ASCII rows, or a short binary body. ## Raises / exceptions The listed structural failures are returned. `Keyword.get/3` raises `FunctionClauseError` when `opts` is not a proper keyword list. An unsupported - `:as` value raises `CaseClauseError`. Malformed row contents may also raise - (`KeyError`, `ArgumentError`, `FunctionClauseError`, or a bitstring match - exception), because scalar/required-property validation is not fully wrapped; - invalid ASCII scalar tokens are currently coerced to `0.0`. + `:as` value raises `CaseClauseError`. Malformed binary row contents may also + raise (`KeyError`, `ArgumentError`, `FunctionClauseError`, or a bitstring match + exception) when scalar unpacking is not fully wrapped. ## Examples @@ -184,8 +191,9 @@ defmodule ExCodecs.Spatial.Codec.PLY do as = Keyword.get(opts, :as, :auto) with {:ok, header, body} <- split_header(data), - {:ok, parsed} <- parse_header(header) do - decode_parsed_body(resolve_as(as, parsed.properties), body, parsed, opts) + {:ok, parsed} <- parse_header(header), + {:ok, as_resolved} <- resolve_as(as, parsed.properties) do + decode_parsed_body(as_resolved, body, parsed, opts) end end @@ -199,15 +207,17 @@ defmodule ExCodecs.Spatial.Codec.PLY do ## Streaming behavior * `source: :file` (or `:auto` path detection): scans the header until - `end_header`, then yields one vertex at a time — binary formats via + `end_header`, then yields one vertex at a time - binary formats via fixed-stride `IO.binread/2`, ASCII via line reads. Peak memory is O(header + one vertex / read chunk), not the whole cloud. * `source: :binary`: still materializes through `decode/2`, then yields. ## Arguments - * `source` (`Path.t() | binary()`) — a path or complete PLY payload. - * `opts` (`keyword()`) — `:source` may be `:auto` (default), `:file`, or + * `source` (`Path.t() | binary()`) - a path or complete PLY payload. + * `opts` (`keyword()`) - `:source` may be `:binary` (default), `:file`, or + `:auto`. Prefer `:file` for paths and `:binary` for payloads; `:auto` is + opt-in path sniffing. PLY also accepts `:as`. `:binary`. Auto mode treats a short path-like binary as a file only when it names a regular file; otherwise it is payload data. `:as` has the same meaning as in `decode/2`. @@ -237,7 +247,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do def stream_decode(source, opts \\ []) def stream_decode(data, opts) when is_binary(data) do - case resolve_ply_source(data, opts) do + case StreamSource.resolve(data, opts, :ply, &path_like_ply?/1) do {:ok, :binary, bin} -> stream_from_decoded_binary(bin, opts) @@ -247,6 +257,9 @@ defmodule ExCodecs.Spatial.Codec.PLY do &next_ply_stream_item/1, &close_ply_stream/1 ) + + {:error, error} -> + Stream.resource(fn -> {:error, error} end, &error_stream/1, fn _ -> :ok end) end end @@ -260,27 +273,8 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end - defp resolve_ply_source(bin, opts) do - case Keyword.get(opts, :source, :auto) do - :binary -> - {:ok, :binary, bin} - - :file -> - {:ok, :file, bin} - - :auto -> - if path_like_ply?(bin) and File.regular?(bin) do - {:ok, :file, bin} - else - {:ok, :binary, bin} - end - end - end - defp path_like_ply?(bin) do - byte_size(bin) < 4096 and not ply_binary?(bin) and - (String.contains?(bin, "/") or String.contains?(bin, "\\") or - String.ends_with?(bin, [".ply", ".PLY"])) + StreamSource.path_like?(bin, &ply_binary?/1, [".ply", ".PLY"]) end @doc false @@ -544,6 +538,37 @@ defmodule ExCodecs.Spatial.Codec.PLY do defp sh_rest_coeff_count(%{sh: [_dc | rest]}), do: length(List.flatten(rest)) defp sh_rest_coeff_count(_), do: 0 + defp validate_comments(comments) when is_list(comments) do + bad = Enum.find(comments, &invalid_comment?/1) + + if bad do + {:error, + Error.new(:invalid_options, + codec: :ply, + message: "PLY comment contains newline or 'end_header': #{inspect(bad)}" + )} + else + :ok + end + end + + defp validate_comments(nil), do: :ok + + defp validate_comments(other) do + {:error, + Error.new(:invalid_options, + codec: :ply, + message: "PLY comments must be a list of strings, got: #{inspect(other)}" + )} + end + + defp invalid_comment?(c) when is_binary(c) do + String.contains?(c, "\n") or String.contains?(c, "\r") or + String.match?(c, ~r/^end_header\b/) + end + + defp invalid_comment?(_), do: true + defp build_header(format, count, props, comments) do format_line = case format do @@ -693,26 +718,34 @@ defmodule ExCodecs.Spatial.Codec.PLY do # Like type_name/1, encode only packs float and uchar values; decode's # unpack/3 still handles the full range of PLY scalar property types. - defp pack(:uchar, v, _), do: <> + defp pack(:uchar, v, _), do: <> defp pack(:float, v, :binary_le), do: <> defp pack(:float, v, :binary_be), do: <> # --- Decode --------------------------------------------------------------- - defp resolve_as(:point_cloud, _), do: :point_cloud - defp resolve_as(:gaussian_cloud, _), do: :gaussian_cloud + defp resolve_as(:point_cloud, _), do: {:ok, :point_cloud} + defp resolve_as(:gaussian_cloud, _), do: {:ok, :gaussian_cloud} defp resolve_as(:auto, props) do names = MapSet.new(Enum.map(props, & &1.name)) if MapSet.member?(names, "f_dc_0") or MapSet.member?(names, "scale_0") or MapSet.member?(names, "rot_0") do - :gaussian_cloud + {:ok, :gaussian_cloud} else - :point_cloud + {:ok, :point_cloud} end end + defp resolve_as(other, _props) do + {:error, + Error.new(:invalid_options, + codec: :ply, + message: "Unsupported :as value: #{inspect(other)} (use :auto, :point_cloud, or :gaussian_cloud)" + )} + end + defp decode_parsed_body(:point_cloud, body, parsed, opts) do with {:ok, points} <- decode_vertices(body, parsed, opts) do {:ok, PointCloud.new(points, metadata: ply_metadata(parsed))} @@ -745,13 +778,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do message: "Expected #{count} vertices, got #{length(lines)}" )} else - points = - Enum.map(lines, fn line -> - values = String.split(line) |> Enum.map(&parse_ascii_number/1) - values_to_point(values, props) - end) - - {:ok, points} + decode_ascii_lines(lines, props) end end @@ -771,13 +798,52 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end + defp decode_ascii_lines(lines, props) do + Enum.reduce_while(lines, {:ok, []}, fn line, {:ok, acc} -> + case decode_ascii_line(line, props) do + {:ok, point} -> {:cont, {:ok, [point | acc]}} + {:error, _} = err -> {:halt, err} + end + end) + |> case do + {:ok, points} -> {:ok, Enum.reverse(points)} + {:error, _} = err -> err + end + end + + defp decode_ascii_line(line, props) do + with {:ok, values} <- parse_ascii_tokens(String.split(line)) do + values_to_point(values, props) + end + end + + defp parse_ascii_tokens(tokens) do + Enum.reduce_while(tokens, {:ok, []}, fn token, {:ok, acc} -> + case parse_ascii_number(token) do + {:ok, number} -> {:cont, {:ok, [number | acc]}} + :error -> {:halt, {:error, invalid_ascii_number_error(token)}} + end + end) + |> case do + {:ok, values} -> {:ok, Enum.reverse(values)} + {:error, _} = err -> err + end + end + + defp invalid_ascii_number_error(token) do + Error.new(:invalid_data, + codec: :ply, + message: "Invalid ASCII number: #{inspect(token)}" + ) + end + defp decode_vertices_binary(body, props, endian, count, expected, opts) do if accel?(opts) do types = Enum.map(props, & &1.type) case Accel.ply_binary_unpack(body, types, endian, 0, count) do {:ok, {rows, next}} when length(rows) == count and next >= expected -> - {:ok, Enum.map(rows, &values_to_point(&1, props))} + rows_to_points(rows, props) # coveralls-ignore-start {:ok, _} -> @@ -796,13 +862,31 @@ defmodule ExCodecs.Spatial.Codec.PLY do end defp decode_vertices_binary_loop(body, props, endian, count) do - {points, _} = - Enum.map_reduce(1..count, body, fn _, rest -> - {values, next} = unpack_row(rest, props, endian) - {values_to_point(values, props), next} - end) + Enum.reduce_while(1..count, {:ok, {[], body}}, fn _, {:ok, {acc, rest}} -> + {values, next} = unpack_row(rest, props, endian) - {:ok, points} + case values_to_point(values, props) do + {:ok, point} -> {:cont, {:ok, {[point | acc], next}}} + {:error, _} = err -> {:halt, err} + end + end) + |> case do + {:ok, {points, _}} -> {:ok, Enum.reverse(points)} + {:error, _} = err -> err + end + end + + defp rows_to_points(rows, props) do + Enum.reduce_while(rows, {:ok, []}, fn values, {:ok, acc} -> + case values_to_point(values, props) do + {:ok, point} -> {:cont, {:ok, [point | acc]}} + {:error, _} = err -> {:halt, err} + end + end) + |> case do + {:ok, points} -> {:ok, Enum.reverse(points)} + {:error, _} = err -> err + end end defp decode_gaussians(body, parsed, opts) do @@ -833,8 +917,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do Map.get(attrs, "rot_2", 0.0), Map.get(attrs, "rot_3", 0.0) }, - sh: extract_sh(attrs), - metadata: attrs + sh: extract_sh(attrs) ) end @@ -860,51 +943,76 @@ defmodule ExCodecs.Spatial.Codec.PLY do end defp values_to_point(values, props) do - paired = Enum.zip(props, values) |> Map.new(fn {%{name: n}, v} -> {n, v} end) + if length(values) != length(props) do + {:error, + Error.new(:invalid_data, + codec: :ply, + message: "Expected #{length(props)} property values, got #{length(values)}" + )} + else + paired = Enum.zip(props, values) |> Map.new(fn {%{name: n}, v} -> {n, v} end) + build_point_from_paired(paired) + end + end - color = - cond do - Map.has_key?(paired, "alpha") and Map.has_key?(paired, "red") -> - {trunc(paired["red"]), trunc(paired["green"]), trunc(paired["blue"]), - trunc(paired["alpha"])} + defp build_point_from_paired(paired) do + with {:ok, x} <- required_coord(paired, "x"), + {:ok, y} <- required_coord(paired, "y"), + {:ok, z} <- required_coord(paired, "z") do + color = + cond do + Map.has_key?(paired, "alpha") and Map.has_key?(paired, "red") -> + {trunc(paired["red"]), trunc(paired["green"]), trunc(paired["blue"]), + trunc(paired["alpha"])} - Map.has_key?(paired, "red") -> - {trunc(paired["red"]), trunc(paired["green"]), trunc(paired["blue"])} + Map.has_key?(paired, "red") -> + {trunc(paired["red"]), trunc(paired["green"]), trunc(paired["blue"])} - true -> + true -> + nil + end + + normal = + if Map.has_key?(paired, "nx") do + {paired["nx"] * 1.0, paired["ny"] * 1.0, paired["nz"] * 1.0} + else nil - end + end - normal = - if Map.has_key?(paired, "nx") do - {paired["nx"] * 1.0, paired["ny"] * 1.0, paired["nz"] * 1.0} - else - nil - end + known = ["x", "y", "z", "red", "green", "blue", "alpha", "nx", "ny", "nz"] - known = ["x", "y", "z", "red", "green", "blue", "alpha", "nx", "ny", "nz"] + attributes = + paired + |> Enum.reject(fn {key, _value} -> key in known end) + |> Map.new() - attributes = - paired - |> Enum.reject(fn {key, _value} -> key in known end) - |> Map.new() + {:ok, Point.new(x, y, z, color: color, normal: normal, attributes: attributes)} + end + end - Point.new(paired["x"] || 0.0, paired["y"] || 0.0, paired["z"] || 0.0, - color: color, - normal: normal, - attributes: attributes - ) + defp required_coord(paired, name) do + case Map.fetch(paired, name) do + {:ok, value} -> + {:ok, value} + + :error -> + {:error, + Error.new(:invalid_data, + codec: :ply, + message: "Missing required PLY property #{name}" + )} + end end defp parse_ascii_number(str) do case Integer.parse(str) do {i, ""} -> - i + {:ok, i} _ -> case Float.parse(str) do - {f, _} -> f - :error -> 0.0 + {f, ""} -> {:ok, f} + _ -> :error end end end @@ -946,13 +1054,20 @@ defmodule ExCodecs.Spatial.Codec.PLY do Keyword.get(opts, :ply_format) || Keyword.get(opts, :format, :ascii) end - defp normalize_format(:ascii), do: :ascii - defp normalize_format(:binary), do: :binary_le - defp normalize_format(:binary_le), do: :binary_le - defp normalize_format(:binary_be), do: :binary_be - defp normalize_format(:little), do: :binary_le - defp normalize_format(:big), do: :binary_be - defp normalize_format(_), do: :ascii + defp normalize_format(:ascii), do: {:ok, :ascii} + defp normalize_format(:binary), do: {:ok, :binary_le} + defp normalize_format(:binary_le), do: {:ok, :binary_le} + defp normalize_format(:binary_be), do: {:ok, :binary_be} + defp normalize_format(:little), do: {:ok, :binary_le} + defp normalize_format(:big), do: {:ok, :binary_be} + + defp normalize_format(other) do + {:error, + Error.new(:invalid_options, + codec: :ply, + message: "Unsupported PLY format: #{inspect(other)}" + )} + end # --- Streaming helpers ---------------------------------------------------- @@ -972,8 +1087,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do {:ok, io} -> case scan_ply_header(io, "", 0) do {:ok, parsed, leftover, body_offset} -> - as = resolve_as(Keyword.get(opts, :as, :auto), parsed.properties) - {:ok, init_ply_stream_state(path, io, leftover, parsed, as, opts, body_offset)} + open_ply_stream_resolved(path, io, leftover, parsed, opts, body_offset) {:error, error} -> File.close(io) @@ -985,6 +1099,17 @@ defmodule ExCodecs.Spatial.Codec.PLY do end end + defp open_ply_stream_resolved(path, io, leftover, parsed, opts, body_offset) do + case resolve_as(Keyword.get(opts, :as, :auto), parsed.properties) do + {:ok, as} -> + {:ok, init_ply_stream_state(path, io, leftover, parsed, as, opts, body_offset)} + + {:error, _} = err -> + File.close(io) + err + end + end + defp init_ply_stream_state( _path, io, @@ -1099,16 +1224,18 @@ defmodule ExCodecs.Spatial.Codec.PLY do ], :done_mmap} {:ok, {rows, next_offset}} -> - items = - Enum.map(rows, fn values -> - emit_vertex(values_to_point(values, parsed.properties), as) - end) + case rows_to_stream_items(rows, parsed.properties, as) do + {:ok, items} -> + new_offset = next_offset - body_offset - new_offset = next_offset - body_offset + {items, + {:ok, + {:binary_mmap, ref, types, endian, parsed, as, body_offset, new_offset, + i + length(rows)}}} - {items, - {:ok, - {:binary_mmap, ref, types, endian, parsed, as, body_offset, new_offset, i + length(rows)}}} + {:error, error} -> + {[{:error, error}], :done_mmap} + end # coveralls-ignore-start {:error, _} -> @@ -1133,8 +1260,15 @@ defmodule ExCodecs.Spatial.Codec.PLY do case take_exact_bytes(io, buf, stride) do {:ok, record, rest} -> {values, _} = unpack_row(record, parsed.properties, endian) - item = emit_vertex(values_to_point(values, parsed.properties), as) - {[item], {:ok, {:binary, io, rest, endian, parsed, as, stride, i + 1}}} + + case values_to_point(values, parsed.properties) do + {:ok, point} -> + {[emit_vertex(point, as)], + {:ok, {:binary, io, rest, endian, parsed, as, stride, i + 1}}} + + {:error, error} -> + {[{:error, error}], {:done, io}} + end {:error, error} -> {[{:error, error}], {:done, io}} @@ -1148,9 +1282,13 @@ defmodule ExCodecs.Spatial.Codec.PLY do defp next_ply_stream_item({:ok, {:ascii, io, buf, parsed, as, i}}) do case take_ascii_vertex_line(io, buf) do {:ok, line, rest} -> - values = line |> String.split() |> Enum.map(&parse_ascii_number/1) - item = emit_vertex(values_to_point(values, parsed.properties), as) - {[item], {:ok, {:ascii, io, rest, parsed, as, i + 1}}} + case decode_ascii_line(line, parsed.properties) do + {:ok, point} -> + {[emit_vertex(point, as)], {:ok, {:ascii, io, rest, parsed, as, i + 1}}} + + {:error, error} -> + {[{:error, error}], {:done, io}} + end {:error, error} -> {[{:error, error}], {:done, io}} @@ -1162,6 +1300,19 @@ defmodule ExCodecs.Spatial.Codec.PLY do defp next_ply_stream_item(:done), do: {:halt, :done} defp next_ply_stream_item(:done_mmap), do: {:halt, :done_mmap} + defp rows_to_stream_items(rows, props, as) do + Enum.reduce_while(rows, {:ok, []}, fn values, {:ok, acc} -> + case values_to_point(values, props) do + {:ok, point} -> {:cont, {:ok, [emit_vertex(point, as) | acc]}} + {:error, _} = err -> {:halt, err} + end + end) + |> case do + {:ok, items} -> {:ok, Enum.reverse(items)} + {:error, _} = err -> err + end + end + defp close_ply_stream({:ok, {:binary, io, _, _, _, _, _, _}}), do: File.close(io) defp close_ply_stream({:ok, {:ascii, io, _, _, _, _}}), do: File.close(io) defp close_ply_stream({:ok, {:binary_mmap, _, _, _, _, _, _, _, _}}), do: :ok @@ -1169,7 +1320,7 @@ defmodule ExCodecs.Spatial.Codec.PLY do defp close_ply_stream(:done_mmap), do: :ok defp close_ply_stream(_), do: :ok - defp accel?(opts), do: Keyword.get(opts, :accel, true) != false and Accel.available?() + defp accel?(opts), do: StreamSource.accel?(opts) defp emit_vertex(point, :point_cloud), do: point defp emit_vertex(point, :gaussian_cloud), do: point_to_gaussian(point) diff --git a/lib/ex_codecs/spatial/codec/stream_source.ex b/lib/ex_codecs/spatial/codec/stream_source.ex new file mode 100644 index 0000000..6e38349 --- /dev/null +++ b/lib/ex_codecs/spatial/codec/stream_source.ex @@ -0,0 +1,51 @@ +defmodule ExCodecs.Spatial.Codec.StreamSource do + @moduledoc false + + # Shared `:source` resolution for spatial stream_decode paths. + # Default is `:binary`; `:file` and `:auto` are explicit opt-ins. + + alias ExCodecs.Spatial.Accel + + @type codec :: :spatial_binary | :gsplat | :ply + @type resolved :: {:ok, :binary | :file, binary()} | {:error, ExCodecs.Error.t()} + + @spec resolve(binary(), keyword(), codec(), (binary() -> boolean())) :: resolved() + def resolve(bin, opts, codec, path_like?) + when is_binary(bin) and is_function(path_like?, 1) and + codec in [:spatial_binary, :gsplat, :ply] do + case Keyword.get(opts, :source, :binary) do + :binary -> + {:ok, :binary, bin} + + :file -> + {:ok, :file, bin} + + :auto -> + if path_like?.(bin) and File.regular?(bin) do + {:ok, :file, bin} + else + {:ok, :binary, bin} + end + + other -> + {:error, + ExCodecs.Error.new(:invalid_options, + codec: codec, + message: "Unsupported :source #{inspect(other)}; use :binary, :file, or :auto" + )} + end + end + + @spec path_like?(binary(), (binary() -> boolean()), [String.t()]) :: boolean() + def path_like?(bin, magic?, extensions) + when is_binary(bin) and is_function(magic?, 1) and is_list(extensions) do + byte_size(bin) < 4096 and not magic?.(bin) and + (String.contains?(bin, "/") or String.contains?(bin, "\\") or + String.ends_with?(bin, extensions)) + end + + @spec accel?(keyword()) :: boolean() + def accel?(opts) when is_list(opts) do + Keyword.get(opts, :accel, true) != false and Accel.available?() + end +end diff --git a/lib/ex_codecs/spatial/gaussian.ex b/lib/ex_codecs/spatial/gaussian.ex index 6e4e183..c2941f8 100644 --- a/lib/ex_codecs/spatial/gaussian.ex +++ b/lib/ex_codecs/spatial/gaussian.ex @@ -118,17 +118,35 @@ defmodule ExCodecs.Spatial.Gaussian do """ @spec new({number(), number(), number()}, keyword()) :: t() def new({x, y, z}, opts \\ []) when is_number(x) and is_number(y) and is_number(z) do + sh = Keyword.get(opts, :sh) + :ok = validate_sh(sh) + %__MODULE__{ position: {x * 1.0, y * 1.0, z * 1.0}, rotation: float4(Keyword.get(opts, :rotation, {1.0, 0.0, 0.0, 0.0})), scale: float3(Keyword.get(opts, :scale, {1.0, 1.0, 1.0})), opacity: Keyword.get(opts, :opacity, 1.0) * 1.0, color: float3(Keyword.get(opts, :color, {0.5, 0.5, 0.5})), - sh: Keyword.get(opts, :sh), + sh: sh, metadata: Keyword.get(opts, :metadata, %{}) } end + defp validate_sh(nil), do: :ok + defp validate_sh([]), do: :ok + + defp validate_sh([[_ | _] | _] = groups) do + if Enum.all?(groups, &is_list/1) do + :ok + else + raise ArgumentError, "sh must be a list of lists of numbers, got: #{inspect(groups)}" + end + end + + defp validate_sh(other) do + raise ArgumentError, "sh must be nil, [], or a list of lists of numbers, got: #{inspect(other)}" + end + defp float3({a, b, c}) when is_number(a) and is_number(b) and is_number(c) do {a * 1.0, b * 1.0, c * 1.0} end diff --git a/lib/ex_codecs/spatial/point.ex b/lib/ex_codecs/spatial/point.ex index 1cba240..ded245a 100644 --- a/lib/ex_codecs/spatial/point.ex +++ b/lib/ex_codecs/spatial/point.ex @@ -51,6 +51,10 @@ defmodule ExCodecs.Spatial.Point do User attributes keyed by strings, with numeric or binary values. For example, `%{"intensity" => 0.82, "classification" => 2}`. Atom keys passed to `new/4` are normalized to strings. + + PLY encode (the only codec that stores attributes) accepts **numeric** + attribute values only; binary values will cause an encode error. Use + cloud-level `Metadata` for non-numeric payloads. """ @type attributes :: %{optional(String.t()) => number() | binary()} @@ -96,6 +100,13 @@ defmodule ExCodecs.Spatial.Point do * `FunctionClauseError` if a coordinate is not numeric, or if `opts` is not a keyword list accepted by `Keyword.get/3`. + ## Notes + + Literal `%Point{}` construction with atom-keyed attributes bypasses + stringification; codecs that read attributes (e.g. PLY) treat atom keys as + missing and default them to `0.0`. Always use `Point.new/4` or pre-stringify + keys when the point will be encoded. + ## Examples iex> p = ExCodecs.Spatial.Point.new(1.0, 2.0, 3.0) diff --git a/lib/ex_codecs/spatial/point_cloud.ex b/lib/ex_codecs/spatial/point_cloud.ex index 506ce26..4837cdb 100644 --- a/lib/ex_codecs/spatial/point_cloud.ex +++ b/lib/ex_codecs/spatial/point_cloud.ex @@ -229,7 +229,9 @@ defmodule ExCodecs.Spatial.PointCloud do end @doc """ - Appends many points (via repeated `add_point/2`). + Appends many points in one concat and one bounds fold (O(n)). + + Prefer this over repeated `add_point/2` for bulk inserts. ## Arguments @@ -253,7 +255,32 @@ defmodule ExCodecs.Spatial.PointCloud do 2 """ @spec add_points(t(), [Point.t()]) :: t() - def add_points(%__MODULE__{} = cloud, new_points) when is_list(new_points) do - Enum.reduce(new_points, cloud, fn p, acc -> add_point(acc, p) end) + def add_points(%__MODULE__{points: points, bounds: bounds} = cloud, []) do + %{cloud | points: points, bounds: bounds} + end + + def add_points(%__MODULE__{points: points, bounds: bounds} = cloud, new_points) + when is_list(new_points) do + new_bounds = + case bounds do + nil -> + Bounds.from_points(new_points) + + %Bounds{} = b -> + Enum.reduce(new_points, b, fn %Point{} = point, acc -> + {x, y, z} = Point.coords(point) + + %Bounds{ + min_x: min(acc.min_x, x), + min_y: min(acc.min_y, y), + min_z: min(acc.min_z, z), + max_x: max(acc.max_x, x), + max_y: max(acc.max_y, y), + max_z: max(acc.max_z, z) + } + end) + end + + %{cloud | points: points ++ new_points, bounds: new_bounds} end end diff --git a/lib/ex_codecs/spatial/stream.ex b/lib/ex_codecs/spatial/stream.ex index 4d286f1..9dabc25 100644 --- a/lib/ex_codecs/spatial/stream.ex +++ b/lib/ex_codecs/spatial/stream.ex @@ -6,7 +6,7 @@ defmodule ExCodecs.Spatial.Stream do * **EXCP** (`:spatial_binary`), **GSPL** (`:gsplat`), and **PLY** with `source: :file` (or `:auto` path detection) read the header, then one - record/vertex at a time from disk — O(header + one record) memory. + record/vertex at a time from disk - O(header + one record) memory. Binary PLY uses a fixed stride; ASCII PLY reads lines. * In-memory binaries still **materialize** through `decode/2`, then yield. A future Rust backend may memory-map large binaries. @@ -15,12 +15,13 @@ defmodule ExCodecs.Spatial.Stream do encoding still collects in memory first. Prefer `source: :file` when the argument is a filesystem path, or - `source: :binary` when it is an encoded payload. With `:auto` (default), a - binary is treated as a path only when it looks path-like (under 4 KiB, no - `ply`/`EXCP`/`GSPL` magic prefix, and contains `/` or `\\` or ends with - `.ply`/`.excp`/`.gspl`/`.bin`) **and** `File.regular?/1` is true. A real - file without separators/extensions is not auto-opened; a short slash-containing - binary that happens to be a regular path may be misread as a file. + `source: :binary` when it is an encoded payload (`:binary` is the default). + With explicit `:auto`, a short path-like binary that names a regular file is + opened; otherwise the argument is treated as payload bytes. Prefer explicit + `:file` / `:binary` for untrusted input - `:auto` can open local paths. + Path-like means under 4 KiB, no `ply`/`EXCP`/`GSPL` magic prefix, and + contains `/` or `\\` or ends with `.ply`/`.excp`/`.gspl`/`.bin`, **and** + `File.regular?/1` is true. """ alias ExCodecs.Error @@ -35,10 +36,11 @@ defmodule ExCodecs.Spatial.Stream do ## Arguments - * `source` (`Path.t() | binary()`) — a path or complete encoded payload. + * `source` (`Path.t() | binary()`) - a path or complete encoded payload. Paths and payloads are both binaries, so `:source` controls resolution. - * `opts` (`keyword()`) — requires `:format`: `:ply`, - `:spatial_binary`, or `:gsplat`. `:source` may be `:auto` (default), + * `opts` (`keyword()`) - requires `:format`: `:ply`, + `:spatial_binary`, or `:gsplat`. `:source` may be `:binary` (default), + `:file`, or `:auto`. `:file`, or `:binary`; PLY also accepts `:as`. ## Returns @@ -48,10 +50,10 @@ defmodule ExCodecs.Spatial.Stream do enumeration and represented by exactly one `{:error, %ExCodecs.Error{}}` element: - * `reason: :invalid_options` — `:format` is absent. - * `reason: :unsupported_codec` — `:format` is unknown. - * `reason: :io_error` — a selected file cannot be opened or read. - * `reason: :invalid_data` — the payload is malformed, unsupported, or + * `reason: :invalid_options` - `:format` is absent. + * `reason: :unsupported_codec` - `:format` is unknown. + * `reason: :io_error` - a selected file cannot be opened or read. + * `reason: :invalid_data` - the payload is malformed, unsupported, or truncated. ## Raises / exceptions @@ -113,9 +115,9 @@ defmodule ExCodecs.Spatial.Stream do ## Arguments - * `enumerable` (`Enumerable.t()`) — `%Point{}` or `%Gaussian{}` elements. + * `enumerable` (`Enumerable.t()`) - `%Point{}` or `%Gaussian{}` elements. The list should be homogeneous and contain valid struct shapes. - * `opts` (`keyword()`) — requires `:format`: `:ply`, + * `opts` (`keyword()`) - requires `:format`: `:ply`, `:spatial_binary`, or `:gsplat`. Remaining keys are forwarded to `ExCodecs.Spatial.encode/2`. @@ -200,11 +202,11 @@ defmodule ExCodecs.Spatial.Stream do ## Arguments - * `data` (`PointCloud.t() | GaussianCloud.t() | Enumerable.t()`) — a cloud, + * `data` (`PointCloud.t() | GaussianCloud.t() | Enumerable.t()`) - a cloud, or an enumerable of `%Point{}`/`%Gaussian{}` values. - * `path` (`Path.t()`) — destination path; parent directories must already + * `path` (`Path.t()`) - destination path; parent directories must already exist. - * `opts` (`keyword()`) — the options for `ExCodecs.Spatial.encode/2`, with + * `opts` (`keyword()`) - the options for `ExCodecs.Spatial.encode/2`, with `:format` required for enumerable input and defaulting to `:ply` for an already-built cloud. For streaming EXCP/GSPL writes, pass `:schema` (see codec docs). diff --git a/lib/ex_codecs/spatial/transform.ex b/lib/ex_codecs/spatial/transform.ex index df00efa..c02cdf7 100644 --- a/lib/ex_codecs/spatial/transform.ex +++ b/lib/ex_codecs/spatial/transform.ex @@ -2,8 +2,10 @@ defmodule ExCodecs.Spatial.Transform do @moduledoc """ Rigid/similarity transform metadata, not applied to points by this library. - Codecs may embed transforms in headers when a format supports it. Application - of transforms is left to higher-level code. + **Reserved for future use.** No shipped codec reads or writes `Transform` + fields today. The module is provided so that formats which embed pose + metadata (e.g. EXCP v2) can use a shared type when that support lands. + Application of transforms is left to higher-level code. ## Fields diff --git a/livebooks/01_introduction.livemd b/livebooks/01_introduction.livemd index a0cfab7..745cc99 100644 --- a/livebooks/01_introduction.livemd +++ b/livebooks/01_introduction.livemd @@ -40,21 +40,19 @@ decompressing. ExCodecs is one framework with category APIs shaped for their data: binary registry codecs use `ExCodecs.encode/3` / `decode/3`, while point-cloud and Gaussian formats use `ExCodecs.Spatial`. +**Binary registry interface:** + ```elixir -# Binary registry interface: -# {:ok, encoded} = ExCodecs.encode(:some_codec, data) -# {:ok, decoded} = ExCodecs.decode(:some_codec, encoded) -# decoded == data # round-trip guarantee -# -# Spatial category interface: -# {:ok, encoded} = ExCodecs.Spatial.encode(cloud, format: :ply) -# {:ok, decoded} = ExCodecs.Spatial.decode(encoded, format: :ply) +{:ok, encoded} = ExCodecs.encode(:some_codec, data) +{:ok, decoded} = ExCodecs.decode(:some_codec, encoded) +decoded == data # round-trip guarantee ``` - +**Spatial category interface:** -``` -nil +```elixir +{:ok, encoded} = ExCodecs.Spatial.encode(cloud, format: :ply) +{:ok, decoded} = ExCodecs.Spatial.decode(encoded, format: :ply) ``` ## Why Codecs Matter @@ -65,9 +63,9 @@ Compression and encoding are fundamental to production systems: | --------------------- | ------------------------------------------- | | **Storage cost** | Compressed data uses less disk space | | **Network bandwidth** | Smaller payloads mean faster transfers | -| **Data integrity** | Decode failures detect corruption | +| **Data integrity** | Some framed codecs detect corruption on decode; raw LZ4/Snappy blocks may not | | **Memory efficiency** | Compressed caches fit more in RAM | -| **Scientific data** | Blosc2 shuffle+compress slashes array sizes | +| **Scientific data** | Blosc2 shuffle+compress reduces typed-array size | A codec framework gives you a **single consistent API** over multiple algorithms, so you can choose the right tool per workload without rewriting integration code. @@ -75,13 +73,11 @@ A codec framework gives you a **single consistent API** over multiple algorithms ExCodecs is a **codec framework**, not just a compression library: -1. **Specialized category APIs** — one framework without unsafe argument overloading -2. **Shared catalog** — query binary and spatial codecs, support, and metadata -3. **Extensible** — binary codecs use `ExCodecs.Codec`; domain categories can define suitable contracts -4. **NIF-native** — Rust-powered NIFs for production throughput -5. **Consistent errors** — structured `%ExCodecs.Error{}` for all failure modes - -Let's see it in action. +1. **Specialized category APIs** - one framework without unsafe argument overloading +2. **Shared catalog** - query binary and spatial codecs, support, and metadata +3. **Extensible** - binary codecs use `ExCodecs.Codec`; domain categories can define suitable contracts +4. **NIF-native** - Rust NIFs for compression throughput +5. **Consistent errors** - structured `%ExCodecs.Error{}` for all failure modes ```elixir original = String.duplicate(~S""" diff --git a/livebooks/03_codec_comparison.livemd b/livebooks/03_codec_comparison.livemd index bcb8598..90cc7ad 100644 --- a/livebooks/03_codec_comparison.livemd +++ b/livebooks/03_codec_comparison.livemd @@ -270,7 +270,7 @@ size (e.g. bzip2 `block_size` × ~100 KiB; see livebook 02). ```elixir profile_data = %{ - "Speed King" => %{ + "Low latency" => %{ best: [:lz4, :snappy], why: "Fastest encode/decode, ideal for hot paths, caching, and real-time systems" }, @@ -284,7 +284,7 @@ profile_data = %{ }, "Numeric Arrays" => %{ best: [:blosc2], - why: "Shuffle+compress slashes size of typed data. Threaded for large arrays" + why: "Shuffle+compress reduces size of typed data (single-threaded NIF)" } } @@ -308,9 +308,9 @@ end ## Numeric Arrays Codecs: [:blosc2] - Shuffle+compress slashes size of typed data. Threaded for large arrays + Shuffle+compress reduces size of typed data (single-threaded NIF) -## Speed King +## Low latency Codecs: [:lz4, :snappy] Fastest encode/decode, ideal for hot paths, caching, and real-time systems @@ -393,58 +393,18 @@ Suggested options: [cname: :zstd, clevel: 5, shuffle: :byte, typesize: 8] ## Decision Flowchart -```elixir -flowchart = """ -When choosing a codec, follow this decision path: - -1. Is your data typed numerical arrays? - YES → Use Blosc2 (with appropriate shuffle and typesize) - NO → Continue - -2. Is latency critical (hot path, real-time)? - YES → Use LZ4 (fastest) or Snappy (low overhead) - NO → Continue - -3. Is storage cost the primary concern? - YES → Use Bzip2 (best ratio) or Zstd with high level - NO → Continue - -4. Default choice: - → Use Zstd (level 3) - → Good ratio, fast decompression, configurable -""" - -IO.puts(flowchart) -``` - - - -``` When choosing a codec, follow this decision path: 1. Is your data typed numerical arrays? - YES → Use Blosc2 (with appropriate shuffle and typesize) - NO → Continue - + - YES → Use Blosc2 (with appropriate shuffle and typesize) + - NO → Continue 2. Is latency critical (hot path, real-time)? - YES → Use LZ4 (fastest) or Snappy (low overhead) - NO → Continue - + - YES → Use LZ4 (fastest) or Snappy (low overhead) + - NO → Continue 3. Is storage cost the primary concern? - YES → Use Bzip2 (best ratio) or Zstd with high level - NO → Continue - -4. Default choice: - → Use Zstd (level 3) - → Good ratio, fast decompression, configurable - -``` - - - -``` -:ok -``` + - YES → Use Bzip2 (best ratio) or Zstd with high level + - NO → Continue +4. Default choice: Zstd (level 3) - good ratio, fast decompression, configurable ## Codec Feature Matrix diff --git a/livebooks/04_building_storage_systems.livemd b/livebooks/04_building_storage_systems.livemd index 4970dfc..6198cab 100644 --- a/livebooks/04_building_storage_systems.livemd +++ b/livebooks/04_building_storage_systems.livemd @@ -166,7 +166,10 @@ end ### Using the Compressed Store ```elixir -{:ok, _pid} = CompressedStore.start_link(codec: :zstd, opts: [level: 3]) +case CompressedStore.start_link(codec: :zstd, opts: [level: 3]) do + {:ok, _pid} -> :ok + {:error, {:already_started, _pid}} -> :ok +end CompressedStore.put(:doc1, String.duplicate("Hello, World! ", 1000)) CompressedStore.put(:doc2, :crypto.strong_rand_bytes(4096) |> Base.encode64()) @@ -274,7 +277,10 @@ end ### Using the Multi-Codec Store ```elixir -{:ok, _pid} = MultiCodecStore.start_link() +case MultiCodecStore.start_link() do + {:ok, _pid} -> :ok + {:error, {:already_started, _pid}} -> :ok +end datasets = %{ repetitive: String.duplicate("abcabcabc", 5000), @@ -489,6 +495,24 @@ end {:ok, <<40, 181, 47, 253, 32, 9, 73, 0, 0, 115, 111, 109, 101, 32, 100, 97, 116, 97>>} ``` +```elixir +# Exercise the error branches with real failures. +results = + [ + {:ok, _} = ExCodecs.encode(:zstd, "ok"), + {:error, %ExCodecs.Error{reason: :invalid_options}} = + ExCodecs.encode(:zstd, "x", level: 99), + {:error, %ExCodecs.Error{reason: :decompression_failed}} = + ExCodecs.decode(:zstd, "not a zstd frame"), + {:error, %ExCodecs.Error{reason: :unsupported_codec}} = + ExCodecs.encode(:nonexistent, "x") + ] + +for {tag, {status, _} = result} <- Enum.zip([:ok, :bad_level, :bad_frame, :bad_codec], results) do + IO.puts("#{String.pad_trailing("#{tag}", 12)} -> #{inspect(result)}") +end +``` + ```elixir # Bound decompression for untrusted payloads (default is 256 MiB). # Classic bomb shape: 64 KiB of zeros → tiny compressed input. @@ -557,7 +581,7 @@ defmodule CompressWithFallback do |> Enum.reduce_while({:error, :no_codec_available}, fn codec, _acc -> case ExCodecs.encode(codec, data) do {:ok, compressed} -> {:halt, {:ok, {codec, compressed}}} - {:error, _} -> {:cont, {:error, :codec_failed}} + {:error, err} -> {:cont, {:error, {codec, err.reason}}} end end) end @@ -589,6 +613,7 @@ defmodule CompressedETS do ETS-backed storage with transparent compression. """ def new(name \\ :compressed_store) do + if :ets.whereis(name) != :undefined, do: :ets.delete(name) :ets.new(name, [:set, :public, :named_table]) name end @@ -704,7 +729,14 @@ defmodule CompressedFileIO do end defp bin_to_codec_atom(bin) do - bin |> String.trim_trailing() |> String.to_atom() + case bin |> String.trim_trailing() do + "zstd" -> :zstd + "lz4" -> :lz4 + "snappy" -> :snappy + "bzip2" -> :bzip2 + "blosc2" -> :blosc2 + other -> {:error, {:unknown_codec, other}} + end end end ``` diff --git a/livebooks/05_zarr_style_workloads.livemd b/livebooks/05_zarr_style_workloads.livemd index e7920ca..69f4190 100644 --- a/livebooks/05_zarr_style_workloads.livemd +++ b/livebooks/05_zarr_style_workloads.livemd @@ -3,22 +3,11 @@ ```elixir # Use this install to work with the source code # Mix.install( -# [ -# {:ex_codecs, path: Path.join(__DIR__, "..")}, -# {:rustler, "~> 0.36"}, -# {:jason, "~> 1.4"}, -# {:kino, "~> 0.14"}, -# {:kino_vega_lite, "~> 0.1.13"} -# ], -# config: [rustler_precompiled: [force_build: [ex_codecs: true]]] +# [{:ex_codecs, path: Path.join(__DIR__, "..")}], +# config: [rustler_precompiled: [force_build: [ex_codecs: true]]] # ) -Mix.install( [ - {:ex_codecs, "~> 0.2.3"}, - {:jason, "~> 1.4"}, - {:kino, "~> 0.14"}, - {:kino_vega_lite, "~> 0.1.13"} - ]) +Mix.install([{:ex_codecs, "~> 0.2.3"}]) ``` ## Series @@ -34,13 +23,8 @@ Mix.install( [ ## Scientific Dataset Compression -ExCodecs is designed to serve as the codec foundation for scientific computing -libraries like ExZarr and ExArrow. This notebook demonstrates compression -patterns optimized for numerical and scientific data. - -```elixir -alias ExCodecs.Compression -``` +This notebook demonstrates compression patterns that array stores such as +[Zarr](https://zarr.dev/) use, optimized for numerical and scientific data. ## Array Data Layout @@ -57,7 +41,7 @@ Blosc2 was specifically designed for this workload: # Generate a synthetic dataset (64-bit float array) n_elements = 100_000 data = :binary.copy(<<3.14159265::float-64>>, n_elements) -byte_size = byte_size(data) +data_size = byte_size(data) # Compare codecs on regular numeric data codecs = [:zstd, :lz4, :snappy, :bzip2, :blosc2] @@ -65,7 +49,7 @@ codecs = [:zstd, :lz4, :snappy, :bzip2, :blosc2] results = for codec <- codecs do {time, {:ok, compressed}} = :timer.tc(fn -> ExCodecs.encode(codec, data) end) - ratio = Float.round(byte_size / byte_size(compressed), 2) + ratio = Float.round(data_size / byte_size(compressed), 2) {codec, byte_size: byte_size(compressed), ratio: ratio, time_us: time} end @@ -101,14 +85,10 @@ Large datasets are typically split into chunks, each compressed independently: ```elixir chunk_size = 1024 * 8 # 8 KiB chunks -large_data = :crypto.strong_rand_bytes(1024 * 1024) # 1 MiB +large_data = :binary.copy(<<3.14159265::float-64>>, 131_072) # 1 MiB of compressible floats chunks = - large_data - |> binary_part(0, byte_size(large_data)) - |> then(fn data -> - for <>, do: chunk - end) + for <>, do: chunk compressed_chunks = chunks diff --git a/mix.exs b/mix.exs index 477f48b..c1aa336 100644 --- a/mix.exs +++ b/mix.exs @@ -107,9 +107,27 @@ defmodule ExCodecs.MixProject do source_url: @source_url, source_ref: "v#{@version}", extras: - Path.wildcard("guides/**/*.md") ++ - Path.wildcard("livebooks/**/*.livemd") ++ - ["docs/spatial_formats.md"], + [ + "guides/benchmarking_methodology.md", + "guides/choosing_compression_codec.md", + "guides/codec_fundamentals.md", + "guides/compression_fundamentals.md", + "guides/native_architecture.md", + "guides/runtime_codec_discovery.md", + "guides/understanding_blosc2.md", + "guides/understanding_bzip2.md", + "guides/understanding_lz4.md", + "guides/understanding_snappy.md", + "guides/understanding_spatial_codecs.md", + "guides/understanding_zstd.md", + "livebooks/01_introduction.livemd", + "livebooks/02_compression_fundamentals.livemd", + "livebooks/03_codec_comparison.livemd", + "livebooks/04_building_storage_systems.livemd", + "livebooks/05_zarr_style_workloads.livemd", + "livebooks/06_spatial_codecs.livemd", + "docs/spatial_formats.md" + ], groups_for_extras: [ Guides: ~r"guides/", Livebooks: ~r"livebooks/", diff --git a/native/ex_codecs_native/src/atoms.rs b/native/ex_codecs_native/src/atoms.rs index 68c9caf..8354b7d 100644 --- a/native/ex_codecs_native/src/atoms.rs +++ b/native/ex_codecs_native/src/atoms.rs @@ -11,4 +11,5 @@ atoms! { decompression_failed, nif_not_loaded, output_limit_exceeded, + io_error, } diff --git a/native/ex_codecs_native/src/blosc2_codec.rs b/native/ex_codecs_native/src/blosc2_codec.rs index a44bfbd..23dab40 100644 --- a/native/ex_codecs_native/src/blosc2_codec.rs +++ b/native/ex_codecs_native/src/blosc2_codec.rs @@ -56,7 +56,9 @@ pub fn blosc2_compress<'a>( }; let filter = filter_from_shuffle(shuffle.clamp(0, 2) as u8); - if data.len() > i32::MAX as usize { + let overhead = BLOSC2_MAX_OVERHEAD as usize; + let max_src = (i32::MAX as usize).saturating_sub(overhead); + if data.len() > max_src { return err(env, atoms::invalid_data()); } @@ -76,8 +78,14 @@ pub fn blosc2_compress<'a>( }; let src = data.as_slice(); - let srcsize = src.len() as i32; - let mut destsize = (src.len() + BLOSC2_MAX_OVERHEAD) as i32; + let Ok(srcsize) = i32::try_from(src.len()) else { + blosc2_free_ctx(ctx); + return err(env, atoms::invalid_data()); + }; + let Ok(mut destsize) = i32::try_from(src.len() + overhead) else { + blosc2_free_ctx(ctx); + return err(env, atoms::invalid_data()); + }; if destsize < BLOSC2_MAX_OVERHEAD as i32 { destsize = BLOSC2_MAX_OVERHEAD as i32; } @@ -91,7 +99,13 @@ pub fn blosc2_compress<'a>( } if n == 0 { // Destination too small — retry with a larger buffer. - let destsize2 = ((src.len() * 2) + BLOSC2_MAX_OVERHEAD).max(BLOSC2_MAX_OVERHEAD) as i32; + let Some(destsize2_usize) = src.len().checked_mul(2).map(|n| n + overhead) else { + return err(env, atoms::compression_failed()); + }; + let destsize2_usize = destsize2_usize.max(overhead); + let Ok(destsize2) = i32::try_from(destsize2_usize) else { + return err(env, atoms::compression_failed()); + }; let mut dest2 = vec![0u8; destsize2 as usize]; let cparams2 = CParams { compcode, diff --git a/native/ex_codecs_native/src/spatial.rs b/native/ex_codecs_native/src/spatial.rs index 539b057..77f71a3 100644 --- a/native/ex_codecs_native/src/spatial.rs +++ b/native/ex_codecs_native/src/spatial.rs @@ -539,14 +539,16 @@ fn ply_binary_unpack<'a>( ) } -#[rustler::nif] +#[rustler::nif(schedule = "DirtyIo")] fn spatial_mmap_open<'a>(env: Env<'a>, path: String) -> Term<'a> { + // Caller must not truncate/replace the file while the mmap resource is live + // (SIGBUS can kill the BEAM). Prefer plain IO for untrusted paths. match File::open(&path).and_then(|f| unsafe { Mmap::map(&f) }) { Ok(mmap) => { let resource = ResourceArc::new(MappedSpatial { mmap }); (atoms::ok(), resource).encode(env) } - Err(_) => err(env, atoms::invalid_data()), + Err(_) => err(env, atoms::io_error()), } } @@ -609,13 +611,13 @@ fn ply_binary_unpack_mmap<'a>( } /// Append packed body bytes to a file (chunked encode helper). -#[rustler::nif(schedule = "DirtyCpu")] +#[rustler::nif(schedule = "DirtyIo")] fn spatial_append_file<'a>(env: Env<'a>, path: String, data: Binary) -> Term<'a> { match File::options().append(true).create(true).open(&path) { Ok(mut file) => match file.write_all(data.as_slice()) { Ok(()) => atoms::ok().encode(env), - Err(_) => err(env, atoms::invalid_data()), + Err(_) => err(env, atoms::io_error()), }, - Err(_) => err(env, atoms::invalid_data()), + Err(_) => err(env, atoms::io_error()), } } diff --git a/native/ex_codecs_native/src/zstd_codec.rs b/native/ex_codecs_native/src/zstd_codec.rs index 3142ca6..23b833a 100644 --- a/native/ex_codecs_native/src/zstd_codec.rs +++ b/native/ex_codecs_native/src/zstd_codec.rs @@ -23,7 +23,8 @@ pub fn zstd_compress<'a>(env: Env<'a>, data: Binary, level: i32) -> Term<'a> { #[rustler::nif(schedule = "DirtyCpu")] pub fn zstd_decompress<'a>(env: Env<'a>, data: Binary, max_output_size: u64) -> Term<'a> { - match structured_zstd::decoding::StreamingDecoder::new(data.as_slice()) { + let mut remaining = data.as_slice(); + match structured_zstd::decoding::StreamingDecoder::new(&mut remaining) { Ok(mut decoder) => { let mut out = Vec::new(); let mut buf = [0u8; 8192]; @@ -39,6 +40,10 @@ pub fn zstd_decompress<'a>(env: Env<'a>, data: Binary, max_output_size: u64) -> Err(_) => return err(env, atoms::decompression_failed()), } } + // Single-frame API: reject concatenated frames / trailing garbage. + if !remaining.is_empty() { + return err(env, atoms::decompression_failed()); + } ok_binary(env, &out) } Err(_) => err(env, atoms::decompression_failed()), diff --git a/test/ex_codecs/codec_registry_test.exs b/test/ex_codecs/codec_registry_test.exs index 836235b..8a1d2b1 100644 --- a/test/ex_codecs/codec_registry_test.exs +++ b/test/ex_codecs/codec_registry_test.exs @@ -83,7 +83,8 @@ defmodule ExCodecs.CodecRegistryTest do end test "filters available codecs by category" do - assert CodecRegistry.available_codecs(:spatial) == [:gsplat, :ply, :spatial_binary] + spatial = CodecRegistry.available_codecs(:spatial) + assert Enum.all?([:gsplat, :ply, :spatial_binary], &(&1 in spatial)) assert :zstd in CodecRegistry.available_codecs(:compression) refute :ply in CodecRegistry.available_codecs(:compression) end @@ -149,7 +150,8 @@ defmodule ExCodecs.CodecRegistryTest do test "returns shared-catalog spatial entries" do codecs = CodecRegistry.codecs_by_category(:spatial) - assert Enum.map(codecs, & &1.name) == [:gsplat, :ply, :spatial_binary] + names = Enum.map(codecs, & &1.name) + assert Enum.all?([:gsplat, :ply, :spatial_binary], &(&1 in names)) assert Enum.all?(codecs, &(&1.interface == :spatial)) end end diff --git a/test/ex_codecs/codec_test.exs b/test/ex_codecs/codec_test.exs index 7f888ac..ce1c9ff 100644 --- a/test/ex_codecs/codec_test.exs +++ b/test/ex_codecs/codec_test.exs @@ -1,39 +1,14 @@ defmodule ExCodecs.CodecTest do use ExUnit.Case, async: true + doctest ExCodecs.Codec + alias ExCodecs.Codec describe "Codec struct" do - test "has required fields" do - codec = %Codec{ - name: :zstd, - category: :compression, - module: ExCodecs.Compression.Zstd, - native?: true, - streaming?: true, - configurable?: true, - version: "1.5.6" - } - - assert codec.name == :zstd - assert codec.category == :compression - assert codec.interface == :binary - assert codec.module == ExCodecs.Compression.Zstd - assert codec.native? == true - assert codec.streaming? == true - assert codec.configurable? == true - assert codec.version == "1.5.6" - end - - test "has default values" do + test "defaults interface to :binary" do codec = %Codec{name: :test} - assert codec.category == nil assert codec.interface == :binary - assert codec.module == nil - assert codec.native? == nil - assert codec.streaming? == nil - assert codec.configurable? == nil - assert codec.version == nil end end diff --git a/test/ex_codecs/compression/blosc2_test.exs b/test/ex_codecs/compression/blosc2_test.exs index 49790d1..923713e 100644 --- a/test/ex_codecs/compression/blosc2_test.exs +++ b/test/ex_codecs/compression/blosc2_test.exs @@ -1,6 +1,9 @@ defmodule ExCodecs.Compression.Blosc2Test do use ExUnit.Case, async: true + @moduletag :doctest + doctest ExCodecs.Compression.Blosc2 + alias ExCodecs.Compression.Blosc2 describe "encode/2" do @@ -141,6 +144,11 @@ defmodule ExCodecs.Compression.Blosc2Test do {:ok, decompressed} = Blosc2.decode(compressed, []) assert decompressed == data end + + test "returns error for invalid magic" do + assert {:error, %ExCodecs.Error{reason: :decompression_failed}} = + Blosc2.decode(<<0xFF, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0>>, []) + end end describe "__codec_info__/0" do diff --git a/test/ex_codecs/compression/bzip2_test.exs b/test/ex_codecs/compression/bzip2_test.exs index 93894fb..7404d86 100644 --- a/test/ex_codecs/compression/bzip2_test.exs +++ b/test/ex_codecs/compression/bzip2_test.exs @@ -1,20 +1,25 @@ defmodule ExCodecs.Compression.Bzip2Test do use ExUnit.Case, async: true + @moduletag :doctest + doctest ExCodecs.Compression.Bzip2 + alias ExCodecs.Compression.Bzip2 describe "encode/2" do - test "compresses data with default block size" do - data = :crypto.strong_rand_bytes(1024) + test "round-trips data at default block size" do + data = String.duplicate("Hello, World! ", 256) assert {:ok, compressed} = Bzip2.encode(data, []) - assert is_binary(compressed) + assert {:ok, ^data} = Bzip2.decode(compressed, []) + assert byte_size(compressed) < byte_size(data) end - test "compresses data with custom block size" do - data = :crypto.strong_rand_bytes(1024) + test "round-trips data at every block size 1..9" do + data = String.duplicate("Hello, World! This is a compression test. ", 256) for bs <- 1..9 do - assert {:ok, _c} = Bzip2.encode(data, block_size: bs) + assert {:ok, compressed} = Bzip2.encode(data, block_size: bs) + assert {:ok, ^data} = Bzip2.decode(compressed, []) end end diff --git a/test/ex_codecs/compression/lz4_test.exs b/test/ex_codecs/compression/lz4_test.exs index 228b048..10ab767 100644 --- a/test/ex_codecs/compression/lz4_test.exs +++ b/test/ex_codecs/compression/lz4_test.exs @@ -1,13 +1,17 @@ defmodule ExCodecs.Compression.Lz4Test do use ExUnit.Case, async: true + @moduletag :doctest + doctest ExCodecs.Compression.Lz4 + alias ExCodecs.Compression.Lz4 describe "encode/2" do - test "compresses data" do - data = :crypto.strong_rand_bytes(1024) + test "round-trips compressible data" do + data = String.duplicate("ab", 4096) assert {:ok, compressed} = Lz4.encode(data, []) - assert is_binary(compressed) + assert {:ok, ^data} = Lz4.decode(compressed, []) + assert byte_size(compressed) < byte_size(data) end test "compresses highly compressible data" do diff --git a/test/ex_codecs/compression/snappy_test.exs b/test/ex_codecs/compression/snappy_test.exs index 680e532..728dd8b 100644 --- a/test/ex_codecs/compression/snappy_test.exs +++ b/test/ex_codecs/compression/snappy_test.exs @@ -1,13 +1,17 @@ defmodule ExCodecs.Compression.SnappyTest do use ExUnit.Case, async: true + @moduletag :doctest + doctest ExCodecs.Compression.Snappy + alias ExCodecs.Compression.Snappy describe "encode/2" do - test "compresses data" do - data = :crypto.strong_rand_bytes(1024) + test "round-trips compressible data" do + data = String.duplicate("hello world ", 1000) assert {:ok, compressed} = Snappy.encode(data, []) - assert is_binary(compressed) + assert {:ok, ^data} = Snappy.decode(compressed, []) + assert byte_size(compressed) < byte_size(data) end test "handles empty data" do diff --git a/test/ex_codecs/compression/version_and_error_test.exs b/test/ex_codecs/compression/version_and_error_test.exs deleted file mode 100644 index 32f67a5..0000000 --- a/test/ex_codecs/compression/version_and_error_test.exs +++ /dev/null @@ -1,163 +0,0 @@ -defmodule ExCodecs.Compression.CodecVersionTest do - use ExUnit.Case, async: true - - alias ExCodecs.Compression.{Blosc2, Bzip2, Lz4, Snappy, Zstd} - alias ExCodecs.Error - - describe "codec_info/0 for all codecs" do - test "zstd returns complete info" do - info = Zstd.__codec_info__() - assert info.name == :zstd - assert info.category == :compression - assert info.module == Zstd - assert info.native? == true - assert info.streaming? == false - assert info.configurable? == true - assert is_binary(info.version) - end - - test "lz4 returns complete info" do - info = Lz4.__codec_info__() - assert info.name == :lz4 - assert info.category == :compression - assert info.module == Lz4 - assert info.native? == true - assert info.configurable? == false - assert is_binary(info.version) - end - - test "snappy returns complete info" do - info = Snappy.__codec_info__() - assert info.name == :snappy - assert info.category == :compression - assert info.module == Snappy - assert info.native? == true - assert info.configurable? == false - assert is_binary(info.version) - end - - test "bzip2 returns complete info" do - info = Bzip2.__codec_info__() - assert info.name == :bzip2 - assert info.category == :compression - assert info.module == Bzip2 - assert info.native? == true - assert info.configurable? == true - assert is_binary(info.version) - end - - test "blosc2 returns complete info" do - info = Blosc2.__codec_info__() - assert info.name == :blosc2 - assert info.category == :compression - assert info.module == Blosc2 - assert info.native? == true - assert info.streaming? == false - assert info.configurable? == true - assert is_binary(info.version) - end - end - - describe "compression error paths" do - test "zstd returns error for corrupt data" do - assert {:error, _} = ExCodecs.decode(:zstd, <<0, 1, 2, 3>>) - end - - test "lz4 returns error for corrupt data" do - assert {:error, _} = ExCodecs.decode(:lz4, <<0, 1, 2, 3>>) - end - - test "snappy returns error for corrupt data" do - assert {:error, _} = ExCodecs.decode(:snappy, <<0, 1, 2, 3>>) - end - - test "bzip2 returns error for corrupt data" do - assert {:error, _} = ExCodecs.decode(:bzip2, <<0, 1, 2, 3>>) - end - - test "blosc2 returns error for corrupt data" do - assert {:error, _} = ExCodecs.decode(:blosc2, <<0, 1, 2, 3>>) - end - - test "blosc2 returns error for invalid magic" do - assert {:error, %Error{reason: reason}} = - ExCodecs.decode(:blosc2, <<0xFF, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0>>) - - assert reason in [:invalid_data, :decompression_failed] - end - - test "blosc2 validates cname option" do - data = :crypto.strong_rand_bytes(1024) - - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:blosc2, data, cname: :invalid) - end - - test "blosc2 validates clevel option" do - data = :crypto.strong_rand_bytes(1024) - - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:blosc2, data, clevel: 10) - end - - test "blosc2 validates shuffle option" do - data = :crypto.strong_rand_bytes(1024) - - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:blosc2, data, shuffle: :invalid) - end - - test "blosc2 accepts bit shuffle" do - floats = for i <- 1..2048, into: <<>>, do: <> - - assert {:ok, _} = ExCodecs.encode(:blosc2, floats, shuffle: :bit) - end - - test "blosc2 validates typesize option" do - data = :crypto.strong_rand_bytes(1024) - - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:blosc2, data, typesize: 0) - end - - test "zstd invalid level returns error" do - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:zstd, "data", level: 0) - - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:zstd, "data", level: 23) - end - - test "bzip2 invalid block_size returns error" do - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:bzip2, "data", block_size: 0) - - assert {:error, %Error{reason: :invalid_options}} = - ExCodecs.encode(:bzip2, "data", block_size: 10) - end - end - - describe "compression category" do - test "Compression.available_codecs returns list" do - codecs = ExCodecs.Compression.available_codecs() - assert is_list(codecs) - assert length(codecs) >= 5 - assert Enum.any?(codecs, &(&1.name == :zstd)) - end - - test "Compression.compress delegates to encode" do - data = String.duplicate("AAAA", 256) - {:ok, c1} = ExCodecs.Compression.compress(:zstd, data) - {:ok, c2} = ExCodecs.encode(:zstd, data) - assert c1 == c2 - end - - test "Compression.decompress delegates to decode" do - data = :crypto.strong_rand_bytes(1024) - {:ok, compressed} = ExCodecs.encode(:zstd, data) - {:ok, d1} = ExCodecs.Compression.decompress(:zstd, compressed) - {:ok, d2} = ExCodecs.decode(:zstd, compressed) - assert d1 == d2 - end - end -end diff --git a/test/ex_codecs/compression/wire_interop_test.exs b/test/ex_codecs/compression/wire_interop_test.exs new file mode 100644 index 0000000..2e927e4 --- /dev/null +++ b/test/ex_codecs/compression/wire_interop_test.exs @@ -0,0 +1,43 @@ +defmodule ExCodecs.Compression.WireInteropTest do + use ExUnit.Case, async: true + + @moduletag :wire_interop + + @fixture_root Path.expand("../../fixtures", __DIR__) + + describe "zstd CLI golden frame" do + test "decompresses level3.bin from zstd(1)" do + src = File.read!(Path.join(@fixture_root, "zstd/level3.src")) + bin = File.read!(Path.join(@fixture_root, "zstd/level3.bin")) + + assert {:ok, ^src} = ExCodecs.decode(:zstd, bin) + end + end + + describe "bzip2 CLI golden stream" do + test "decompresses default.bin from bzip2(1)" do + src = File.read!(Path.join(@fixture_root, "bzip2/default.src")) + bin = File.read!(Path.join(@fixture_root, "bzip2/default.bin")) + + assert {:ok, ^src} = ExCodecs.decode(:bzip2, bin) + end + end + + describe "lz4_flex size-prepended golden" do + test "decompresses size_prepended.bin from lz4_flex" do + src = File.read!(Path.join(@fixture_root, "lz4/size_prepended.src")) + bin = File.read!(Path.join(@fixture_root, "lz4/size_prepended.bin")) + + assert {:ok, ^src} = ExCodecs.decode(:lz4, bin) + end + end + + describe "snap raw golden" do + test "decompresses raw.bin from snap::raw" do + src = File.read!(Path.join(@fixture_root, "snappy/raw.src")) + bin = File.read!(Path.join(@fixture_root, "snappy/raw.bin")) + + assert {:ok, ^src} = ExCodecs.decode(:snappy, bin) + end + end +end diff --git a/test/ex_codecs/compression/zstd_test.exs b/test/ex_codecs/compression/zstd_test.exs index 50c1b7f..0af91eb 100644 --- a/test/ex_codecs/compression/zstd_test.exs +++ b/test/ex_codecs/compression/zstd_test.exs @@ -1,6 +1,9 @@ defmodule ExCodecs.Compression.ZstdTest do use ExUnit.Case, async: true + @moduletag :doctest + doctest ExCodecs.Compression.Zstd + alias ExCodecs.Compression.Zstd describe "encode/2" do @@ -11,11 +14,13 @@ defmodule ExCodecs.Compression.ZstdTest do assert byte_size(compressed) < byte_size(data) end - test "compresses data with custom level" do - data = :crypto.strong_rand_bytes(1024) + test "round-trips data at every level 1..22" do + data = String.duplicate("The quick brown fox jumps over the lazy dog. ", 64) for level <- 1..22 do - assert {:ok, _compressed} = Zstd.encode(data, level: level) + assert {:ok, compressed} = Zstd.encode(data, level: level) + assert {:ok, ^data} = Zstd.decode(compressed, []) + assert byte_size(compressed) < byte_size(data) end end @@ -131,15 +136,4 @@ defmodule ExCodecs.Compression.ZstdTest do end end - describe "round-trip property" do - test "zstd round-trip preserves random data" do - for _ <- 1..50 do - size = :rand.uniform(10_000) - data = :crypto.strong_rand_bytes(size) - {:ok, compressed} = Zstd.encode(data, []) - {:ok, decompressed} = Zstd.decode(compressed, []) - assert decompressed == data - end - end - end end diff --git a/test/ex_codecs/coverage_boost_test.exs b/test/ex_codecs/coverage_boost_test.exs index ceac9c4..c6101ff 100644 --- a/test/ex_codecs/coverage_boost_test.exs +++ b/test/ex_codecs/coverage_boost_test.exs @@ -177,13 +177,19 @@ defmodule ExCodecs.CoverageBoostTest do g = struct(Gaussian, position: {0.0, 0.0, 0.0}, + color: {0.1, 0.2, 0.3}, sh: [0.1, 0.2, 0.3, 0.4, 0.5, 0.6] ) cloud = GaussianCloud.new([g]) - assert {:ok, bin} = Gsplat.encode(cloud) - assert {:ok, decoded} = Gsplat.decode(bin) - assert hd(decoded.gaussians).sh != nil + assert {:ok, bin} = Gsplat.encode(cloud, accel: false) + # Flat SH must drop 3 DC floats → sh_rest=3 (not 5 from dropping only the head). + assert byte_size(bin) == 18 + 56 + 12 + assert {:ok, decoded} = Gsplat.decode(bin, accel: false) + assert [[_, _, _], [r1, r2, r3]] = hd(decoded.gaussians).sh + assert_in_delta r1, 0.4, 1.0e-6 + assert_in_delta r2, 0.5, 1.0e-6 + assert_in_delta r3, 0.6, 1.0e-6 # Valid header with one gaussian requiring SH floats that are missing. header = <<"GSPL", 1::little-16, 0::little-16, 1::little-64, 3::little-16>> @@ -200,7 +206,9 @@ defmodule ExCodecs.CoverageBoostTest do assert {:ok, _} = PLY.encode(cloud, format: :binary_le) assert {:ok, _} = PLY.encode(cloud, format: :little) assert {:ok, _} = PLY.encode(cloud, format: :big) - assert {:ok, _} = PLY.encode(cloud, format: :unknown_defaults_ascii) + + assert {:error, %{reason: :invalid_options}} = + PLY.encode(cloud, format: :unknown_defaults_ascii) assert {:error, %{reason: :invalid_data}} = PLY.decode(123) @@ -346,7 +354,7 @@ defmodule ExCodecs.CoverageBoostTest do assert {:ok, _} = PLY.encode(sparse, format: :ascii) - assert {:ok, weird_cloud} = + assert {:error, %{reason: :invalid_data}} = PLY.decode(""" ply format ascii 1.0 @@ -359,8 +367,6 @@ defmodule ExCodecs.CoverageBoostTest do 1 2 3 not-a-number """) - assert hd(weird_cloud.points).attributes["weird"] == 0.0 - assert {:ok, _} = PLY.decode("ply\nformat ascii 1.0\nelement vertex 0\nend_header") @@ -429,7 +435,8 @@ defmodule ExCodecs.CoverageBoostTest do File.write!(path, excp) assert [%Point{}] = - Spatial.stream_decode(path, format: :spatial_binary) |> Enum.to_list() + Spatial.stream_decode(path, format: :spatial_binary, source: :file) + |> Enum.to_list() end end end diff --git a/test/ex_codecs/documentation_test.exs b/test/ex_codecs/documentation_test.exs index c631844..840c33c 100644 --- a/test/ex_codecs/documentation_test.exs +++ b/test/ex_codecs/documentation_test.exs @@ -1,6 +1,9 @@ defmodule ExCodecs.DocumentationTest do use ExUnit.Case, async: true + # Heading/contract linter for moduledocs — exclude with `mix test --exclude lint` + @moduletag :lint + @documented_modules [ ExCodecs, ExCodecs.Application, diff --git a/test/ex_codecs/nif_test.exs b/test/ex_codecs/nif_test.exs index 023daa3..44ae151 100644 --- a/test/ex_codecs/nif_test.exs +++ b/test/ex_codecs/nif_test.exs @@ -34,7 +34,7 @@ defmodule ExCodecs.NIFTest do end test "wraps unknown error atoms" do - assert {:error, %ExCodecs.Error{reason: :some_unknown_error, codec: :snappy}} = + assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :snappy, details: :some_unknown_error}} = NIF.wrap(:snappy, {:error, :some_unknown_error}) end end diff --git a/test/ex_codecs/property_expansion_test.exs b/test/ex_codecs/property_expansion_test.exs new file mode 100644 index 0000000..4d12bf7 --- /dev/null +++ b/test/ex_codecs/property_expansion_test.exs @@ -0,0 +1,149 @@ +defmodule ExCodecs.PropertyExpansionTest do + use ExUnit.Case, async: true + use ExUnitProperties + + alias ExCodecs.Error + alias ExCodecs.Spatial.{Bounds, Point} + + @moduletag timeout: 60_000 + + property "zstd round-trips across small levels" do + check all( + data <- binary(min_length: 0, max_length: 512), + level <- integer(1..5), + max_runs: 25 + ) do + {:ok, compressed} = ExCodecs.encode(:zstd, data, level: level) + {:ok, decompressed} = ExCodecs.decode(:zstd, compressed) + assert decompressed == data + end + end + + property "bzip2 round-trips across block sizes" do + check all( + data <- binary(min_length: 1, max_length: 256), + block_size <- integer(1..9), + max_runs: 20 + ) do + {:ok, compressed} = ExCodecs.encode(:bzip2, data, block_size: block_size) + {:ok, decompressed} = ExCodecs.decode(:bzip2, compressed) + assert decompressed == data + end + end + + property "blosc2 round-trips across clevel with typesize-aligned data" do + check all( + typesize <- integer(1..8), + n_elems <- integer(1..16), + clevel <- integer(0..9), + cname <- member_of([:lz4, :zstd]), + shuffle <- member_of([:none, :byte]), + max_runs: 30 + ) do + data = :binary.copy(<<7>>, typesize * n_elems) + + {:ok, compressed} = + ExCodecs.encode(:blosc2, data, + cname: cname, + clevel: clevel, + shuffle: shuffle, + typesize: typesize + ) + + {:ok, decompressed} = ExCodecs.decode(:blosc2, compressed) + assert decompressed == data + end + end + + property "zstd max_output_size rejects undersized and accepts exact" do + check all(data <- binary(min_length: 2, max_length: 256), max_runs: 25) do + {:ok, compressed} = ExCodecs.encode(:zstd, data) + + assert {:error, %Error{reason: :output_limit_exceeded}} = + ExCodecs.decode(:zstd, compressed, max_output_size: byte_size(data) - 1) + + assert {:ok, ^data} = + ExCodecs.decode(:zstd, compressed, max_output_size: byte_size(data)) + end + end + + property "lz4 max_output_size rejects undersized and accepts exact" do + check all(data <- binary(min_length: 2, max_length: 256), max_runs: 25) do + {:ok, compressed} = ExCodecs.encode(:lz4, data) + + assert {:error, %Error{reason: :output_limit_exceeded}} = + ExCodecs.decode(:lz4, compressed, max_output_size: byte_size(data) - 1) + + assert {:ok, ^data} = + ExCodecs.decode(:lz4, compressed, max_output_size: byte_size(data)) + end + end + + property "truncation never crashes framed codecs" do + check all( + codec <- member_of([:zstd, :bzip2, :blosc2]), + data <- binary(min_length: 8, max_length: 128), + max_runs: 40 + ) do + encode_opts = + case codec do + :blosc2 -> [shuffle: :none, typesize: 1] + _ -> [] + end + + {:ok, compressed} = ExCodecs.encode(codec, data, encode_opts) + truncated = corrupt(compressed, :truncate) + + assert {:error, %Error{}} = ExCodecs.decode(codec, truncated) + end + end + + property "magic-byte corruption never crashes zstd or bzip2" do + check all( + codec <- member_of([:zstd, :bzip2]), + data <- binary(min_length: 8, max_length: 128), + max_runs: 40 + ) do + {:ok, compressed} = ExCodecs.encode(codec, data) + corrupted = corrupt(compressed, :flip_magic) + + assert {:error, %Error{}} = ExCodecs.decode(codec, corrupted) + end + end + + property "Bounds.from_points contains every input point" do + check all( + points <- list_of(point_gen(), min_length: 1, max_length: 20), + max_runs: 40 + ) do + bounds = Bounds.from_points(points) + refute is_nil(bounds) + + for %Point{x: x, y: y, z: z} <- points do + assert Bounds.contains?(bounds, {x, y, z}) + end + end + end + + defp point_gen do + gen all( + x <- float(min: -1000.0, max: 1000.0), + y <- float(min: -1000.0, max: 1000.0), + z <- float(min: -1000.0, max: 1000.0) + ) do + Point.new(x, y, z) + end + end + + defp corrupt(bin, :truncate) when byte_size(bin) > 1 do + binary_part(bin, 0, byte_size(bin) - 1) + end + + defp corrupt(_bin, :truncate), do: <<>> + + # Flip the frame magic / leading byte so parsers reject without relying on + # optional content checksums (zstd may accept mid-frame bit flips). + defp corrupt(<>, :flip_magic) do + <> + end +end diff --git a/test/ex_codecs/property_test.exs b/test/ex_codecs/property_test.exs index 34457b0..9ca9c5d 100644 --- a/test/ex_codecs/property_test.exs +++ b/test/ex_codecs/property_test.exs @@ -3,7 +3,7 @@ defmodule ExCodecs.PropertyTest do use ExUnitProperties property "zstd round-trip preserves data" do - check all(data <- binary(min: 0, max: 10_000)) do + check all(data <- binary(min_length: 0, max_length: 10_000)) do {:ok, compressed} = ExCodecs.encode(:zstd, data) {:ok, decompressed} = ExCodecs.decode(:zstd, compressed) assert decompressed == data @@ -11,7 +11,7 @@ defmodule ExCodecs.PropertyTest do end property "lz4 round-trip preserves data" do - check all(data <- binary(min: 1, max: 10_000)) do + check all(data <- binary(min_length: 1, max_length: 10_000)) do {:ok, compressed} = ExCodecs.encode(:lz4, data) {:ok, decompressed} = ExCodecs.decode(:lz4, compressed) assert decompressed == data @@ -19,7 +19,7 @@ defmodule ExCodecs.PropertyTest do end property "snappy round-trip preserves data" do - check all(data <- binary(min: 1, max: 10_000)) do + check all(data <- binary(min_length: 1, max_length: 10_000)) do {:ok, compressed} = ExCodecs.encode(:snappy, data) {:ok, decompressed} = ExCodecs.decode(:snappy, compressed) assert decompressed == data @@ -27,7 +27,7 @@ defmodule ExCodecs.PropertyTest do end property "bzip2 round-trip preserves data" do - check all(data <- binary(min: 1, max: 10_000)) do + check all(data <- binary(min_length: 1, max_length: 10_000)) do {:ok, compressed} = ExCodecs.encode(:bzip2, data) {:ok, decompressed} = ExCodecs.decode(:bzip2, compressed) assert decompressed == data @@ -35,8 +35,7 @@ defmodule ExCodecs.PropertyTest do end property "blosc2 round-trip preserves data (random)" do - check all(size <- integer(8..10_000)) do - data = :crypto.strong_rand_bytes(size) + check all(data <- binary(min_length: 8, max_length: 10_000)) do {:ok, compressed} = ExCodecs.encode(:blosc2, data) {:ok, decompressed} = ExCodecs.decode(:blosc2, compressed) assert decompressed == data @@ -45,7 +44,7 @@ defmodule ExCodecs.PropertyTest do property "compression reduces size for repeated data" do check all( - byte <- string(Enum.take(?a..?z, 1), min_length: 1, max_length: 1), + byte <- string(?a..?z, length: 1), count <- integer(100..10_000) ) do data = String.duplicate(byte, count) @@ -117,17 +116,8 @@ defmodule ExCodecs.PropertyTest do end end - property "codec_info is consistent for all codecs" do - check all( - codec <- - one_of([ - constant(:zstd), - constant(:lz4), - constant(:snappy), - constant(:bzip2), - constant(:blosc2) - ]) - ) do + test "codec_info is consistent for all codecs" do + for codec <- [:zstd, :lz4, :snappy, :bzip2, :blosc2] do {:ok, info} = ExCodecs.codec_info(codec) assert info.name == codec assert info.category == :compression diff --git a/test/ex_codecs/spatial/accel_coverage_test.exs b/test/ex_codecs/spatial/accel_coverage_test.exs index e27174b..3415a78 100644 --- a/test/ex_codecs/spatial/accel_coverage_test.exs +++ b/test/ex_codecs/spatial/accel_coverage_test.exs @@ -5,7 +5,7 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do alias ExCodecs.Spatial.Codec.{Binary, Gsplat, PLY} alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} - unless ExCodecs.Spatial.Accel.available?(), + unless Accel.available?(), do: @moduletag(skip: "spatial Accel NIF not loaded") defp tmp(ext) do @@ -19,8 +19,17 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do test "chunk_size, ply_type_tag, pack/unpack, mmap, and append_file" do assert Accel.chunk_size() == 4096 - for t <- [:char, :uchar, :short, :ushort, :int, :uint, :float, :double] do - assert is_integer(Accel.ply_type_tag(t)) + for {t, tag} <- [ + char: 1, + uchar: 2, + short: 3, + ushort: 4, + int: 5, + uint: 6, + float: 7, + double: 8 + ] do + assert Accel.ply_type_tag(t) == tag end points = [ @@ -56,7 +65,7 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do assert_in_delta z, 3.0, 1.0e-5 assert {:ok, {[[_, _, _]], _}} = - Accel.ply_binary_unpack(ply_body, [:float, :float, :float], true, 0, 1) + Accel.ply_binary_unpack(ply_body, [:float, :float, :float], :little, 0, 1) path = tmp(".excp") on_exit(fn -> File.rm(path) end) @@ -122,7 +131,7 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do assert {:error, _} = Accel.gspl_pack([bad_g], 0) end - test "row helpers cover nil SH and nested SH" do + test "row helpers cover nil SH, nested SH, and flat SH (drops 3 DC)" do p = Point.new(1, 2, 3, color: {1, 2, 3, 4}, normal: nil) assert {1.0, 2.0, 3.0, {1, 2, 3, 4}, nil} = Accel.point_to_row(p) assert %Point{} = Accel.row_to_point(Accel.point_to_row(p)) @@ -134,6 +143,16 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do g1 = Gaussian.new({0, 0, 0}, color: {0.1, 0.2, 0.3}, sh: [[0.1, 0.2, 0.3], [1.0, 2.0, 3.0]]) row1 = Accel.gaussian_to_row(g1, 3) assert %Gaussian{} = Accel.row_to_gaussian(row1, 3) + + g_flat = + struct(Gaussian, + position: {0.0, 0.0, 0.0}, + color: {0.1, 0.2, 0.3}, + sh: [0.1, 0.2, 0.3, 0.4, 0.5, 0.6] + ) + + {_, _, _, _, _, flat_rest} = Accel.gaussian_to_row(g_flat, 3) + assert flat_rest == [0.4, 0.5, 0.6] end end @@ -305,14 +324,32 @@ defmodule ExCodecs.Spatial.AccelCoverageTest do ) ]) - # Empty SH list hits sh_rest_values/1 catch-all list clause (non-empty lists - # match the [_dc | rest] head clause first). + # Empty SH list hits sh_rest_values/1 empty-list clause. assert {:ok, _} = Gsplat.encode( GaussianCloud.new([struct(Gaussian, position: {0.0, 0.0, 0.0}, sh: [])]), accel: false ) + # Flat SH drops the first 3 DC floats (not just the head element). + flat = + GaussianCloud.new([ + struct(Gaussian, + position: {0.0, 0.0, 0.0}, + color: {0.1, 0.2, 0.3}, + sh: [0.1, 0.2, 0.3, 0.4, 0.5, 0.6] + ) + ]) + + assert {:ok, flat_bin} = Gsplat.encode(flat, accel: false) + # 18-byte header + 14 f32 base + 3 f32 SH rest + assert byte_size(flat_bin) == 18 + 56 + 12 + assert {:ok, flat_decoded} = Gsplat.decode(flat_bin, accel: false) + assert [[_, _, _], rest] = hd(flat_decoded.gaussians).sh + assert_in_delta hd(rest), 0.4, 1.0e-6 + assert_in_delta Enum.at(rest, 1), 0.5, 1.0e-6 + assert_in_delta Enum.at(rest, 2), 0.6, 1.0e-6 + assert {:ok, bin} = Gsplat.encode(cloud, accel: false) assert {:ok, decoded} = Gsplat.decode(bin, accel: false) assert length(decoded.gaussians) == 2 diff --git a/test/ex_codecs/spatial/bounds_test.exs b/test/ex_codecs/spatial/bounds_test.exs index 1c8c699..463dfa0 100644 --- a/test/ex_codecs/spatial/bounds_test.exs +++ b/test/ex_codecs/spatial/bounds_test.exs @@ -1,6 +1,8 @@ defmodule ExCodecs.Spatial.BoundsTest do use ExUnit.Case, async: true + doctest ExCodecs.Spatial.Bounds + alias ExCodecs.Spatial.{Bounds, Point} test "computes bounds from points" do diff --git a/test/ex_codecs/spatial/codec/binary_test.exs b/test/ex_codecs/spatial/codec/binary_test.exs index fae9929..3cb6e17 100644 --- a/test/ex_codecs/spatial/codec/binary_test.exs +++ b/test/ex_codecs/spatial/codec/binary_test.exs @@ -53,9 +53,9 @@ defmodule ExCodecs.Spatial.Codec.BinaryTest do assert length(points) == 2 assert hd(points).color == {10, 20, 30} - # :auto path detection + # :auto path detection (opt-in; default :source is :binary) assert [%Point{}, %Point{}] = - Binary.stream_decode(path) |> Enum.to_list() + Binary.stream_decode(path, source: :auto) |> Enum.to_list() end test "stream_decode from truncated file yields invalid_data" do diff --git a/test/ex_codecs/spatial/coverage_test.exs b/test/ex_codecs/spatial/coverage_test.exs index 086d643..145ea3c 100644 --- a/test/ex_codecs/spatial/coverage_test.exs +++ b/test/ex_codecs/spatial/coverage_test.exs @@ -15,10 +15,13 @@ defmodule ExCodecs.Spatial.CoverageTest do end test "unsupported version and truncation" do - assert {:error, _} = Binary.decode(<<"EXCP", 99::little-16, 0::little-16, 1::little-64>>) - assert {:error, _} = Binary.decode(<<"EXCP", 1::little-16, 1::little-16, 1::little-64, 0, 1>>) + assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :spatial_binary}} = + Binary.decode(<<"EXCP", 99::little-16, 0::little-16, 1::little-64>>) - assert [{:error, _}] = + assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :spatial_binary}} = + Binary.decode(<<"EXCP", 1::little-16, 1::little-16, 1::little-64, 0, 1>>) + + assert [{:error, %ExCodecs.Error{}}] = Binary.stream_decode(<<"nope">>) |> Enum.to_list() end @@ -29,16 +32,18 @@ defmodule ExCodecs.Spatial.CoverageTest do describe "gsplat edges" do test "unsupported version and truncation" do - assert {:error, _} = + assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :gsplat}} = Gsplat.decode(<<"GSPL", 9::little-16, 0::little-16, 1::little-64, 0::little-16>>) - assert {:error, _} = + assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :gsplat}} = Gsplat.decode( <<"GSPL", 1::little-16, 0::little-16, 1::little-64, 0::little-16, 1, 2>> ) - assert [{:error, _}] = Gsplat.stream_decode(<<"nope">>) |> Enum.to_list() - assert {:error, _} = Gsplat.encode(%{}) + assert [{:error, %ExCodecs.Error{}}] = Gsplat.stream_decode(<<"nope">>) |> Enum.to_list() + + assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :gsplat}} = + Gsplat.encode(%{}) end end @@ -116,10 +121,10 @@ defmodule ExCodecs.Spatial.CoverageTest do end test "missing magic / format / vertex element" do - assert {:error, _} = PLY.decode("format ascii 1.0\nend_header\n") - assert {:error, _} = PLY.decode("ply\nelement vertex 0\nend_header\n") + assert {:error, %{reason: :invalid_data}} = PLY.decode("format ascii 1.0\nend_header\n") + assert {:error, %{reason: :invalid_data}} = PLY.decode("ply\nelement vertex 0\nend_header\n") - assert {:error, _} = + assert {:error, %{reason: :invalid_data}} = PLY.decode("ply\nformat ascii 1.0\nend_header\n") end @@ -132,7 +137,7 @@ defmodule ExCodecs.Spatial.CoverageTest do end_header """ - assert {:error, _} = PLY.decode(data) + assert {:error, %{reason: :invalid_data}} = PLY.decode(data) end test "stream_decode file path" do @@ -143,7 +148,7 @@ defmodule ExCodecs.Spatial.CoverageTest do {:ok, bin} = PLY.encode(cloud) File.write!(path, bin) - assert [%Point{x: x}] = PLY.stream_decode(path) |> Enum.to_list() + assert [%Point{x: x}] = PLY.stream_decode(path, source: :file) |> Enum.to_list() assert_in_delta x, 9.0, 0.0001 end @@ -168,7 +173,7 @@ defmodule ExCodecs.Spatial.CoverageTest do assert :ok = SpatialStream.encode_to_file([Point.new(1, 2, 3)], path, format: :ply) - assert {:error, _} = + assert {:error, %{reason: :io_error}} = SpatialStream.encode_to_file([Point.new(1, 2, 3)], "/no/such/dir/x.ply", format: :ply ) @@ -184,11 +189,11 @@ defmodule ExCodecs.Spatial.CoverageTest do File.write!(path, bin) assert [%Point{}] = - Spatial.stream_decode(path, format: :spatial_binary) |> Enum.to_list() + Spatial.stream_decode(path, format: :spatial_binary, source: :file) |> Enum.to_list() end test "stream encode error tuples" do - assert {:error, _} = + assert {:error, %{reason: :invalid_data}} = Spatial.stream_encode([{:error, :boom}], format: :ply) end @@ -223,12 +228,17 @@ defmodule ExCodecs.Spatial.CoverageTest do assert :ok = SpatialStream.encode_to_file(cloud, path, format: :gsplat) assert [%Gaussian{}] = - Spatial.stream_decode(path, format: :gsplat) |> Enum.to_list() + Spatial.stream_decode(path, format: :gsplat, source: :file) |> Enum.to_list() end test "propagates encode errors" do - assert {:error, _} = - SpatialStream.encode_to_file([Point.new(0, 0, 0)], "/tmp/x.gspl", format: :gsplat) + path = + Path.join(System.tmp_dir!(), "ex_codecs_bad_#{System.unique_integer([:positive])}.gspl") + + on_exit(fn -> File.rm(path) end) + + assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :gsplat}} = + SpatialStream.encode_to_file([Point.new(0, 0, 0)], path, format: :gsplat) end end end diff --git a/test/ex_codecs/spatial/stream_test.exs b/test/ex_codecs/spatial/stream_test.exs index d2a44aa..0eab796 100644 --- a/test/ex_codecs/spatial/stream_test.exs +++ b/test/ex_codecs/spatial/stream_test.exs @@ -1,6 +1,8 @@ defmodule ExCodecs.Spatial.StreamTest do use ExUnit.Case, async: true + doctest ExCodecs.Spatial.Stream + alias ExCodecs.Spatial alias ExCodecs.Spatial.{Gaussian, Point, PointCloud} diff --git a/test/ex_codecs/spatial/types_test.exs b/test/ex_codecs/spatial/types_test.exs index 7d821a9..e98eb67 100644 --- a/test/ex_codecs/spatial/types_test.exs +++ b/test/ex_codecs/spatial/types_test.exs @@ -1,6 +1,12 @@ defmodule ExCodecs.Spatial.TypesTest do use ExUnit.Case, async: true + doctest ExCodecs.Spatial.Gaussian + doctest ExCodecs.Spatial.GaussianCloud + doctest ExCodecs.Spatial.Metadata + doctest ExCodecs.Spatial.Transform + doctest ExCodecs.Spatial.PointCloud + alias ExCodecs.Spatial.{ Bounds, Gaussian, diff --git a/test/ex_codecs/spatial_test.exs b/test/ex_codecs/spatial_test.exs index a0dba03..5257170 100644 --- a/test/ex_codecs/spatial_test.exs +++ b/test/ex_codecs/spatial_test.exs @@ -1,6 +1,8 @@ defmodule ExCodecs.SpatialTest do use ExUnit.Case, async: true + doctest ExCodecs.Spatial + alias ExCodecs.Spatial alias ExCodecs.Spatial.{Gaussian, GaussianCloud, Point, PointCloud} @@ -11,16 +13,16 @@ defmodule ExCodecs.SpatialTest do describe "available_formats/0" do test "lists supported formats" do - assert Spatial.available_formats() == [:ply, :spatial_binary, :gsplat] - assert :ply in Spatial.available_formats() - assert :gsplat in Spatial.available_formats() + formats = Spatial.available_formats() + assert Enum.all?([:ply, :spatial_binary, :gsplat], &(&1 in formats)) assert Spatial.supports?(:ply) refute Spatial.supports?(:sog) end test "spatial formats participate in the shared catalog" do assert :ply in ExCodecs.available_codecs() - assert ExCodecs.available_codecs(:spatial) == [:gsplat, :ply, :spatial_binary] + spatial = ExCodecs.available_codecs(:spatial) + assert Enum.all?([:gsplat, :ply, :spatial_binary], &(&1 in spatial)) assert ExCodecs.supports?(:ply) assert {:ok, info} = ExCodecs.codec_info(:ply) diff --git a/test/fixtures/README.md b/test/fixtures/README.md new file mode 100644 index 0000000..071323d --- /dev/null +++ b/test/fixtures/README.md @@ -0,0 +1,24 @@ +# Compression wire-format goldens + +These fixtures pin ExCodecs' pure-Rust backends against reference wire bytes. + +| Codec | Tool | Notes | +|-------|------|--------| +| `zstd/` | `zstd(1)` CLI | Standard Zstandard frame | +| `bzip2/` | `bzip2(1)` CLI | Standard bzip2 stream | +| `lz4/` | `lz4_flex::compress_prepend_size` | Size-prepended block (not LZ4 frame) | +| `snappy/` | `snap::raw` | Raw Snappy block (not framing format) | + +Regenerate (from repo root): + +```sh +# zstd + bzip2 +printf 'payload...' > test/fixtures/zstd/level3.src +zstd -3 -f -o test/fixtures/zstd/level3.bin test/fixtures/zstd/level3.src +cp test/fixtures/zstd/level3.src test/fixtures/bzip2/default.src +bzip2 -c test/fixtures/bzip2/default.src > test/fixtures/bzip2/default.bin + +# lz4 + snappy (uses the same crates as the NIF) +cd native/ex_codecs_native +# one-shot: cargo run --example gen_goldens (see prior generation) +``` diff --git a/test/fixtures/bzip2/default.bin b/test/fixtures/bzip2/default.bin new file mode 100644 index 0000000000000000000000000000000000000000..2964e9c3de4397cb4085cbc740de0c5e43d26e64 GIT binary patch literal 103 zcmV-t0GR(mT4*^jL0KkKS*QVgrvLyWoq#|9f8Yv0Kc${PAOLVs00000s;TNU!$A)~ z(^2Y3nUOMwlG5Vy0wyL#W`#(ckTy3YNZ_IL=yX7|z(l<8TtI;`g-lHfpNgNw+>uTc JBq{)3DZs~JDN+Cc literal 0 HcmV?d00001 diff --git a/test/fixtures/bzip2/default.src b/test/fixtures/bzip2/default.src new file mode 100644 index 0000000..144b10e --- /dev/null +++ b/test/fixtures/bzip2/default.src @@ -0,0 +1 @@ +Hello ExCodecs golden fixture 0123456789 abcdefHello ExCodecs golden fixture 0123456789 abcdefHello ExCodecs golden fixture 0123456789 abcdefHello ExCodecs golden fixture 0123456789 abcdef \ No newline at end of file diff --git a/test/fixtures/lz4/size_prepended.bin b/test/fixtures/lz4/size_prepended.bin new file mode 100644 index 0000000000000000000000000000000000000000..c499ac2c324aa9616d63d52ec38c2154c76a5407 GIT binary patch literal 63 zcmdnPz`*cd!6P*%Ctty}!Z|-BHMv+JJwGQUHBTWev!bN5C{@A0(8$=t)Xdz%QXw%Z OIVCkspP?iH!U6!IClt5< literal 0 HcmV?d00001 diff --git a/test/fixtures/lz4/size_prepended.src b/test/fixtures/lz4/size_prepended.src new file mode 100644 index 0000000..144b10e --- /dev/null +++ b/test/fixtures/lz4/size_prepended.src @@ -0,0 +1 @@ +Hello ExCodecs golden fixture 0123456789 abcdefHello ExCodecs golden fixture 0123456789 abcdefHello ExCodecs golden fixture 0123456789 abcdefHello ExCodecs golden fixture 0123456789 abcdef \ No newline at end of file diff --git a/test/fixtures/snappy/raw.bin b/test/fixtures/snappy/raw.bin new file mode 100644 index 0000000000000000000000000000000000000000..6b97386eb38a0266399ed348f338e07464cca4ab GIT binary patch literal 59 zcmdnPxWgkgCnsOQwZb_+B{jKNAw54QB{feWEwiGev?x` Date: Sat, 18 Jul 2026 15:20:23 -0400 Subject: [PATCH 6/6] fixed formatting --- lib/ex_codecs/nif.ex | 22 ++++++------ lib/ex_codecs/spatial/codec/ply.ex | 3 +- mix.exs | 43 ++++++++++++------------ test/ex_codecs/compression/zstd_test.exs | 1 - test/ex_codecs/nif_test.exs | 3 +- 5 files changed, 37 insertions(+), 35 deletions(-) diff --git a/lib/ex_codecs/nif.ex b/lib/ex_codecs/nif.ex index b41500d..d4779a7 100644 --- a/lib/ex_codecs/nif.ex +++ b/lib/ex_codecs/nif.ex @@ -71,16 +71,18 @@ defmodule ExCodecs.NIF do def wrap(codec, {:error, :invalid_options}), do: {:error, ExCodecs.Error.new(:invalid_options, codec: codec)} - def wrap(codec, {:error, reason}) when is_atom(reason) and reason not in [ - :compression_failed, - :decompression_failed, - :output_limit_exceeded, - :invalid_data, - :invalid_options, - :nif_not_loaded, - :io_error, - :truncated_input - ] do + def wrap(codec, {:error, reason}) + when is_atom(reason) and + reason not in [ + :compression_failed, + :decompression_failed, + :output_limit_exceeded, + :invalid_data, + :invalid_options, + :nif_not_loaded, + :io_error, + :truncated_input + ] do {:error, ExCodecs.Error.new(:invalid_data, codec: codec, diff --git a/lib/ex_codecs/spatial/codec/ply.ex b/lib/ex_codecs/spatial/codec/ply.ex index 7ed510d..f6e1eca 100644 --- a/lib/ex_codecs/spatial/codec/ply.ex +++ b/lib/ex_codecs/spatial/codec/ply.ex @@ -742,7 +742,8 @@ defmodule ExCodecs.Spatial.Codec.PLY do {:error, Error.new(:invalid_options, codec: :ply, - message: "Unsupported :as value: #{inspect(other)} (use :auto, :point_cloud, or :gaussian_cloud)" + message: + "Unsupported :as value: #{inspect(other)} (use :auto, :point_cloud, or :gaussian_cloud)" )} end diff --git a/mix.exs b/mix.exs index c1aa336..b1b09f6 100644 --- a/mix.exs +++ b/mix.exs @@ -106,28 +106,27 @@ defmodule ExCodecs.MixProject do main: "ExCodecs", source_url: @source_url, source_ref: "v#{@version}", - extras: - [ - "guides/benchmarking_methodology.md", - "guides/choosing_compression_codec.md", - "guides/codec_fundamentals.md", - "guides/compression_fundamentals.md", - "guides/native_architecture.md", - "guides/runtime_codec_discovery.md", - "guides/understanding_blosc2.md", - "guides/understanding_bzip2.md", - "guides/understanding_lz4.md", - "guides/understanding_snappy.md", - "guides/understanding_spatial_codecs.md", - "guides/understanding_zstd.md", - "livebooks/01_introduction.livemd", - "livebooks/02_compression_fundamentals.livemd", - "livebooks/03_codec_comparison.livemd", - "livebooks/04_building_storage_systems.livemd", - "livebooks/05_zarr_style_workloads.livemd", - "livebooks/06_spatial_codecs.livemd", - "docs/spatial_formats.md" - ], + extras: [ + "guides/benchmarking_methodology.md", + "guides/choosing_compression_codec.md", + "guides/codec_fundamentals.md", + "guides/compression_fundamentals.md", + "guides/native_architecture.md", + "guides/runtime_codec_discovery.md", + "guides/understanding_blosc2.md", + "guides/understanding_bzip2.md", + "guides/understanding_lz4.md", + "guides/understanding_snappy.md", + "guides/understanding_spatial_codecs.md", + "guides/understanding_zstd.md", + "livebooks/01_introduction.livemd", + "livebooks/02_compression_fundamentals.livemd", + "livebooks/03_codec_comparison.livemd", + "livebooks/04_building_storage_systems.livemd", + "livebooks/05_zarr_style_workloads.livemd", + "livebooks/06_spatial_codecs.livemd", + "docs/spatial_formats.md" + ], groups_for_extras: [ Guides: ~r"guides/", Livebooks: ~r"livebooks/", diff --git a/test/ex_codecs/compression/zstd_test.exs b/test/ex_codecs/compression/zstd_test.exs index 0af91eb..a111257 100644 --- a/test/ex_codecs/compression/zstd_test.exs +++ b/test/ex_codecs/compression/zstd_test.exs @@ -135,5 +135,4 @@ defmodule ExCodecs.Compression.ZstdTest do assert info.configurable? == true end end - end diff --git a/test/ex_codecs/nif_test.exs b/test/ex_codecs/nif_test.exs index 44ae151..9c35c26 100644 --- a/test/ex_codecs/nif_test.exs +++ b/test/ex_codecs/nif_test.exs @@ -34,7 +34,8 @@ defmodule ExCodecs.NIFTest do end test "wraps unknown error atoms" do - assert {:error, %ExCodecs.Error{reason: :invalid_data, codec: :snappy, details: :some_unknown_error}} = + assert {:error, + %ExCodecs.Error{reason: :invalid_data, codec: :snappy, details: :some_unknown_error}} = NIF.wrap(:snappy, {:error, :some_unknown_error}) end end