Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

29 changes: 29 additions & 0 deletions docs/ADMIN_API.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,9 +21,38 @@ body:
```json
{
"dump_url"?: string,
"dump_importer"?: "buffered" | "streaming",
}
```

`dump_url` initializes the new namespace from a SQLite SQL dump (`sqlite3 db .dump` or this
server's `GET /dump` output). Supported schemes are `file:` (an absolute path on the server
host) and `http(s):`. The dump must run inside a transaction and end with `COMMIT`; `ATTACH` is
rejected.

`dump_importer` selects how the dump is loaded (it requires `dump_url`):

- `buffered` (historical): the whole dump is read into memory, parsed, then executed. Memory
usage is proportional to the dump size.
- `streaming`: statements are framed with a resumable port of SQLite's `sqlite3_complete()` state
machine and executed while the dump is still being read. Memory usage is bounded by the
server's queue settings plus the largest single statement (see `--dump-import-*` flags); a
statement larger than `--dump-import-max-statement-size` is rejected with `413`.

When omitted, the server's `--dump-importer` setting (`SQLD_DUMP_IMPORTER`, default `buffered`)
applies. Values are matched case-insensitively; an unknown value is rejected with `422`.

Both importers produce the same data. Known differences:

- the streaming importer stores the schema SQL exactly as written in the dump, whereas the
buffered importer stores the parser's normalized rendering;
- a dump whose *data* contains the word "attach" is rejected by the buffered importer (substring
check) but accepted by the streaming one (statement-level check); a standalone `DETACH`
statement is rejected with `400` by the streaming importer and fails at execution (`500`) with
the buffered one;
- invalid UTF-8 or NUL bytes yield `400` (buffered: `500`), and statements that return rows are
executed with their rows discarded (buffered: `500`).

```HTTP
DELETE /v1/namespaces/:namespace
```
Expand Down
737 changes: 737 additions & 0 deletions docs/STREAMING_DUMP_IMPORT_DESIGN.md

Large diffs are not rendered by default.

26 changes: 26 additions & 0 deletions docs/USER_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -246,6 +246,32 @@ For example, to create a database named `db1`, send the following HTTP request:
curl -X POST http://localhost:8080/v1/namespaces/db1/create
```

### Creating a database from a SQLite dump

A new database can be initialized from a SQL dump produced by `sqlite3 source.db .dump` (or by
this server's `GET /dump` endpoint):

```shell
sqlite3 source.db .dump > /srv/dumps/db1.sql
curl -X POST http://localhost:9090/v1/namespaces/db1/create \
-H 'content-type: application/json' \
-d '{"dump_url": "file:///srv/dumps/db1.sql", "dump_importer": "streaming"}'
```

Two importers are available. `buffered` (the default) reads the whole dump into memory before
executing it. `streaming` executes statements as they arrive and keeps memory usage independent
of the dump size; it is selected per request with `dump_importer` or server-wide with
`--dump-importer streaming` (`SQLD_DUMP_IMPORTER`). The streaming importer is tuned with:

- `--dump-import-max-statement-size` (`SQLD_DUMP_IMPORT_MAX_STATEMENT_SIZE`, default `64MiB`):
a single statement larger than this fails the import with HTTP 413.
- `--dump-import-queue-bytes` (`SQLD_DUMP_IMPORT_QUEUE_BYTES`, default `16MiB`) and
`--dump-import-queue-depth` (`SQLD_DUMP_IMPORT_QUEUE_DEPTH`, default `256`): how many
framed-but-not-yet-executed statements may be in flight.

Either way the dump runs as one transaction, so the WAL and the replication log grow to roughly
the database size before the final `COMMIT`; make sure the disk has room for that.

The name of the database is determined from the `Host` header in the HTTP request.

For example, if you have the following entries in your `/etc/hosts` file:
Expand Down
1 change: 1 addition & 0 deletions libsql-server/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ itertools = "0.10.5"
jsonwebtoken = "9"
libsql = { path = "../libsql/", optional = true }
libsql_replication = { path = "../libsql-replication" }
memchr = "2"
metrics = "0.21.1"
metrics-util = "0.15"
metrics-exporter-prometheus = "0.12.2"
Expand Down
3 changes: 3 additions & 0 deletions libsql-server/src/config.rs
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ use tonic::transport::Channel;
use tower::ServiceExt;

use crate::auth::{Auth, Disabled};
pub use crate::namespace::dump_import::{DumpImportConfig, DumpImporterKind};
use crate::net::{AddrIncoming, Connector};

pub struct RpcClientConfig<C = HttpConnector> {
Expand Down Expand Up @@ -103,6 +104,7 @@ pub struct DbConfig {
pub max_concurrent_requests: u64,
pub disable_intelligent_throttling: bool,
pub connection_creation_timeout: Option<Duration>,
pub dump_import: DumpImportConfig,
}

impl Default for DbConfig {
Expand All @@ -123,6 +125,7 @@ impl Default for DbConfig {
max_concurrent_requests: 128,
disable_intelligent_throttling: false,
connection_creation_timeout: None,
dump_import: DumpImportConfig::default(),
}
}
}
Expand Down
12 changes: 12 additions & 0 deletions libsql-server/src/error.rs
Original file line number Diff line number Diff line change
Expand Up @@ -296,6 +296,16 @@ pub enum LoadDumpError {
NotAFile,
#[error("The passed dump sql is invalid: {0}")]
InvalidSqlInput(String),
#[error("`dump_importer` requires `dump_url`")]
ImporterWithoutDumpUrl,
#[error(
"A dump statement starting at line {line}, column {column} exceeds the maximum allowed size ({limit} bytes)"
)]
StatementTooLarge {
line: u64,
column: usize,
limit: usize,
},
}

impl ResponseError for LoadDumpError {}
Expand All @@ -315,7 +325,9 @@ impl IntoResponse for &LoadDumpError {
| NoCommit
| NotAFile
| DumpFilePathNotAbsolute
| ImporterWithoutDumpUrl
| InvalidSqlInput(_) => self.format_err(StatusCode::BAD_REQUEST),
StatementTooLarge { .. } => self.format_err(StatusCode::PAYLOAD_TOO_LARGE),
}
}
}
Expand Down
26 changes: 21 additions & 5 deletions libsql-server/src/http/admin/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,8 @@ use crate::auth::parse_jwt_keys;
use crate::connection::config::{DatabaseConfig, DurabilityMode};
use crate::error::{Error, LoadDumpError};
use crate::hrana;
use crate::namespace::{DumpStream, NamespaceName, NamespaceStore, RestoreOption};
use crate::namespace::dump_import::DumpImporterKind;
use crate::namespace::{DumpSource, DumpStream, NamespaceName, NamespaceStore, RestoreOption};
use crate::net::Connector;
use crate::LIBSQL_PAGE_SIZE;

Expand Down Expand Up @@ -371,6 +372,10 @@ async fn handle_post_config<C>(
#[derive(Debug, Deserialize)]
struct CreateNamespaceReq {
dump_url: Option<Url>,
/// Which importer loads `dump_url`: `"buffered"` (historical, whole dump in memory) or
/// `"streaming"` (memory-bounded). Defaults to the server's `--dump-importer`.
#[serde(default)]
dump_importer: Option<DumpImporterKind>,
max_db_size: Option<bytesize::ByteSize>,
heartbeat_url: Option<String>,
bottomless_db_id: Option<String>,
Expand Down Expand Up @@ -408,6 +413,10 @@ async fn handle_create_namespace<C: Connector>(
));
}

if req.dump_importer.is_some() && req.dump_url.is_none() {
return Err(LoadDumpError::ImporterWithoutDumpUrl.into());
}

if let Some(ns) = req.shared_schema_name {
if req.shared_schema {
return Err(Error::SharedSchemaCreationError(
Expand All @@ -423,9 +432,10 @@ async fn handle_create_namespace<C: Connector>(
}

let dump = match req.dump_url {
Some(ref url) => {
RestoreOption::Dump(dump_stream_from_url(url, app_state.connector.clone()).await?)
}
Some(ref url) => RestoreOption::Dump(
DumpSource::new(dump_stream_from_url(url, app_state.connector.clone()).await?)
.with_importer(req.dump_importer),
),
None => RestoreOption::Latest,
};

Expand Down Expand Up @@ -474,6 +484,9 @@ async fn handle_fork_namespace<C>(
Ok(())
}

/// Read `file:` dumps in reasonably large chunks; both importers consume the resulting stream.
const DUMP_FILE_READ_CHUNK_SIZE: usize = 64 * 1024;

async fn dump_stream_from_url<C>(url: &Url, connector: C) -> Result<DumpStream, LoadDumpError>
where
C: Connector,
Expand Down Expand Up @@ -507,7 +520,10 @@ where

let f = tokio::fs::File::open(path).await?;

Ok(Box::new(ReaderStream::new(f)))
Ok(Box::new(ReaderStream::with_capacity(
f,
DUMP_FILE_READ_CHUNK_SIZE,
)))
}
scheme => Err(LoadDumpError::UnsupportedUrlScheme(scheme.to_string())),
}
Expand Down
1 change: 1 addition & 0 deletions libsql-server/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -601,6 +601,7 @@ where
encryption_config: self.db_config.encryption_config.clone(),
disable_intelligent_throttling: self.db_config.disable_intelligent_throttling,
connection_creation_timeout: self.db_config.connection_creation_timeout,
dump_import: self.db_config.dump_import.clone(),
};

let (metastore_conn_maker, meta_store_wal_manager) =
Expand Down
36 changes: 34 additions & 2 deletions libsql-server/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -16,8 +16,8 @@ use tracing_subscriber::Layer;
use tracing_subscriber::{prelude::*, EnvFilter};

use libsql_server::config::{
AdminApiConfig, BottomlessConfig, DbConfig, HeartbeatConfig, MetaStoreConfig, RpcClientConfig,
RpcServerConfig, TlsConfig, UserApiConfig,
AdminApiConfig, BottomlessConfig, DbConfig, DumpImportConfig, DumpImporterKind,
HeartbeatConfig, MetaStoreConfig, RpcClientConfig, RpcServerConfig, TlsConfig, UserApiConfig,
};
use libsql_server::net::AddrIncoming;
use libsql_server::version::Version;
Expand Down Expand Up @@ -253,6 +253,28 @@ struct Cli {
#[clap(long, env = "SQLD_CONNECTION_CREATION_TIMEOUT_SEC")]
connection_creation_timeout_sec: Option<u64>,

/// Importer used by `POST /v1/namespaces/:ns/create` with `dump_url` when the request
/// doesn't specify `dump_importer`. `buffered` reads the whole dump into memory;
/// `streaming` executes statements as they arrive with bounded memory.
#[clap(long, env = "SQLD_DUMP_IMPORTER", default_value = "buffered")]
dump_importer: DumpImporterKind,

/// Streaming dump importer: reject any single SQL statement larger than this.
#[clap(
long,
env = "SQLD_DUMP_IMPORT_MAX_STATEMENT_SIZE",
default_value = "64MiB"
)]
dump_import_max_statement_size: ByteSize,

/// Streaming dump importer: maximum bytes of statements framed but not yet executed.
#[clap(long, env = "SQLD_DUMP_IMPORT_QUEUE_BYTES", default_value = "16MiB")]
dump_import_queue_bytes: ByteSize,

/// Streaming dump importer: maximum number of statements framed but not yet executed.
#[clap(long, env = "SQLD_DUMP_IMPORT_QUEUE_DEPTH", default_value = "256")]
dump_import_queue_depth: usize,

/// Allow meta store to recover config from filesystem from older version, if meta store is
/// empty on startup
#[clap(long, env = "SQLD_ALLOW_METASTORE_RECOVERY")]
Expand Down Expand Up @@ -401,6 +423,15 @@ fn make_db_config(config: &Cli) -> anyhow::Result<DbConfig> {
bottomless_replication.encryption_config = encryption_config.clone();
}
}
let dump_import = DumpImportConfig {
default_importer: config.dump_importer,
max_statement_bytes: usize::try_from(config.dump_import_max_statement_size.as_u64())
.context("dump import max statement size doesn't fit in usize")?,
queue_bytes: usize::try_from(config.dump_import_queue_bytes.as_u64())
.context("dump import queue bytes doesn't fit in usize")?,
queue_depth: config.dump_import_queue_depth,
};
dump_import.validate()?;
Ok(DbConfig {
extensions_path: config.extensions_path.clone().map(Into::into),
bottomless_replication,
Expand All @@ -419,6 +450,7 @@ fn make_db_config(config: &Cli) -> anyhow::Result<DbConfig> {
connection_creation_timeout: config
.connection_creation_timeout_sec
.map(|x| Duration::from_secs(x)),
dump_import,
})
}

Expand Down
Loading
Loading