diff --git a/content/docs/search-architecture.es.mdx b/content/docs/search-architecture.es.mdx index 8b25539..b4e1238 100644 --- a/content/docs/search-architecture.es.mdx +++ b/content/docs/search-architecture.es.mdx @@ -1,100 +1,101 @@ --- -title: Arquitectura de búsqueda -description: Cómo se conectan la búsqueda de mensajes, los grafos de conversaciones, los embeddings y las fuentes. +title: Cómo funciona la búsqueda +description: "La parte técnica de la búsqueda de mensajes y la búsqueda por temas: índices, grafo de conversaciones, vectores, orden de resultados, frescura y mediciones." --- -WireCat ofrece dos vías sobre un archivo local compartido: **la búsqueda estricta de mensajes** encuentra mensajes que cumplen una consulta y **la búsqueda de conversaciones** encuentra discusiones por significado y palabras. Los grafos y los embeddings ya existen como etapas opcionales, separadas de la búsqueda Lucene de mensajes. +WireCat tiene dos formas de buscar en un mismo archivo local. **La búsqueda de mensajes** encuentra los mensajes que cumplen una consulta estricta. **La búsqueda por temas** encuentra conversaciones por significado y por palabras en común. Solo comparten el archivo: el grafo de conversaciones y los vectores son etapas opcionales que la búsqueda de mensajes nunca lee. -Esta guía describe la implementación examinada el 4 de octubre de 2026. Los detalles de los comandos están en las guías de [Telegram](./tg/search.md) y [MAX](./max/search.md) y sus páginas del archivo. Los enlaces al código fijan la versión examinada del motor; no implican que todos los CLI instalados ya usen esa versión del SDK. +Esta es la página técnica. Para el uso diario, lee las guías de búsqueda de [Telegram](./tg/search.md) y [MAX](./max/search.md) y sus páginas del archivo. La página describe el motor de [`cli-messaging` 0.148.2](https://github.com/leemour/cli-messaging/tree/v0.148.2), que usan las versiones en desarrollo de tg y max. Las versiones publicadas (tg 0.27, max 0.28) usan 0.147.0, que aún no tiene etiquetas, búsquedas guardadas ni el ajuste de raíces de palabras. -## Elegir la vía adecuada +## Elige la vía adecuada -| Pregunta | Entrada | Resultado | Preparación | +| Pregunta | Comando | Qué devuelve | Qué hace falta antes | | --- | --- | --- | --- | -| «¿Dónde dijo Alice invoice en octubre, con un archivo?» | `messages search` | Mensajes coincidentes, localizadores y contexto cercano opcional | Historial guardado; índice de palabras preparado para condiciones de texto | -| «¿Dónde hablamos de alquilar un piso?» | `conversations search` | Conversaciones ordenadas, fragmentos coincidentes y motivo de recuperación | Conversaciones construidas; embeddings para coincidencias semánticas | -| «¿A qué discusión pertenece este mensaje?» | `conversations show` / `messages links` | Miembros, enlaces candidatos y cadena de padres elegida | Conversaciones construidas | -| «¿Qué se escribió alrededor de este mensaje?» | `messages context` / `--context` | Vecinos cronológicos en su cuenta y chat | Historial guardado; no necesita grafo ni embeddings | +| «¿Dónde dijo Alice factura, en octubre, con un PDF?» | `messages search` | Mensajes que coinciden, sus localizadores y, si quieres, los mensajes vecinos | Historial guardado; índice de palabras listo para consultas con palabras | +| «¿Cuántos mensajes al día mencionan el presupuesto?» | `messages stats` | Recuentos por chat, remitente, día u hora | Lo mismo que la búsqueda de mensajes | +| «¿Dónde hablamos de alquilar un piso?» | `conversations search` | Conversaciones ordenadas, el fragmento que coincide y cómo coincidió | Conversaciones construidas; vectores para coincidir por significado | +| «¿Qué más trató lo mismo?» | `conversations related` | Las conversaciones más cercanas a la de un mensaje | Vectores; no se ejecuta ningún modelo | +| «¿A qué conversación pertenece este mensaje?» | `conversations show` / `messages links` | La conversación y la cadena de padres elegida | Conversaciones construidas | +| «¿Qué había alrededor de este mensaje?» | `messages context` / `--context` en la búsqueda | Vecinos en el tiempo, en la misma cuenta y chat | Solo el historial guardado | -Un tema nativo del mensajero, una conversación inferida y una ventana cronológica son conceptos distintos. Una conversación puede unir mensajes no adyacentes; los mensajes cercanos pueden tratar otro asunto. +Un tema de foro del mensajero, una conversación inferida y una ventana de mensajes vecinos son tres cosas distintas. Una conversación puede unir mensajes lejanos; los mensajes vecinos pueden pertenecer a discusiones distintas. ## 1. Archivo e identidad -Las lecturas, la descarga explícita del historial y la captura de novedades habilitada guardan mensajes en SQLite. La búsqueda lee esa copia y no descarga silenciosamente todo el historial remoto. Los identificadores provider/account/chat/message separan mensajes con el mismo número en fuentes diferentes. Los localizadores completos permiten volver al original. +Las lecturas, las descargas explícitas de historial y la captura en segundo plano guardan los mensajes en un único almacén SQLite compartido. La búsqueda lee esta copia local y nunca descarga historial remoto por su cuenta. Los identificadores de proveedor, cuenta, chat y mensaje separan mensajes con el mismo número de fuentes distintas, y cada resultado lleva un localizador completo para que un agente vuelva al original. -El texto original y los metadatos del proveedor se conservan separados del texto normalizado del índice. El archivo registra rangos de historial y su completitud. Un índice, una conversación construida o un vector no prueban que se haya descargado todo el chat. +El almacén guarda el texto original de cada mensaje y, aparte, su texto normalizado para buscar. También registra qué tramos del historial de cada chat tiene. Un índice, una construcción o un vector no prueban que un chat se descargara entero. -La búsqueda estricta empieza en la cuenta activa. `--source all` o `in:all` amplía explícitamente el ámbito a las cuentas guardadas; un ámbito de proveedor lo reduce. Los nombres se resuelven dentro del ámbito elegido. Si un nombre es ambiguo, hay que usar un identificador preciso o reducir el ámbito. +La búsqueda de mensajes empieza en la cuenta activa. `--source all` o `in:all` la amplía a todas las cuentas del almacén; un proveedor la reduce. Los nombres se resuelven dentro de ese alcance, y un autor o chat ambiguo pide un identificador más preciso. La búsqueda por temas lee una sola cuenta: sus conversaciones construidas, opcionalmente un chat y un límite de tiempo. -La búsqueda de conversaciones está actualmente limitada a una cuenta: sus conversaciones construidas, con filtros opcionales de chat y tiempo. No ofrece un equivalente de `--source all`. La búsqueda de mensajes entre mensajeros no implica búsqueda semántica entre cuentas. +## 2. Búsqueda estricta de mensajes -## 2. Búsqueda estricta de mensajes con Lucene +El proceso: -El flujo de ejecución: - -1. Analiza el texto y crea un AST booleano versionado con posiciones para los errores. Lucene define el lenguaje comprobado de consultas; no significa que haya un servidor Elasticsearch ni un índice Java Lucene. -2. Un registro común valida campos, valores y operadores. El CLI y las peticiones MCP de texto/AST llegan al mismo servicio. -3. Resuelve cuentas, nombres almacenados y límites de fechas. `topic:` requiere un único chat obligatorio porque los IDs de temas son locales al chat. -4. Compila el árbol en condiciones SQLite parametrizadas. Términos y frases usan el índice FTS5 de palabras normalizadas; los metadatos restringen los candidatos. -5. Expande comodines/regex de texto mediante el vocabulario. Las condiciones sobre el cuerpo completo y los detectores preset ejecutan comprobaciones acotadas. Regex usa un autómata limitado en lugar de backtracking sin restricciones. -6. Ordena solo los mensajes que cumplen la consulta y puede recuperar sus vecinos cronológicos en la cuenta/chat original. La relevancia nunca amplía el conjunto booleano. +1. Analizar el texto en un árbol booleano con versión que conserva las posiciones para los errores. La gramática sigue un perfil probado de la sintaxis de consultas de Apache Lucene 9.12.3. Lucene es aquí el lenguaje de consulta, no un servidor de búsqueda ni un índice Java. +2. Validar campos, valores y operadores con un único registro. El texto de la CLI y el texto y los árboles de MCP pasan por el mismo validador y servicio. +3. Resolver el alcance de cuentas, los nombres guardados y los límites de calendario en la zona IANA elegida. `topic:` necesita exactamente un chat, porque los números de tema son locales a cada chat. +4. Compilar el árbol en SQL parametrizado para SQLite. Palabras y frases usan un índice de palabras FTS5 sobre el texto normalizado; los predicados de metadatos reducen los candidatos. +5. Expandir comodines y expresiones regulares de `text:` mediante el vocabulario. Los patrones sobre el texto completo y los presets revisan candidatos dentro de un presupuesto. Las expresiones regulares se ejecutan como un autómata sin retroceso. +6. Ordenar solo los mensajes que coinciden y, si se pide, añadir vecinos en la cuenta y el chat de cada resultado. El orden nunca amplía el resultado. -El motor implementa `text`, `body`, `from`, `chat`, `date`, `kind`, `has`, `topic`, `in` y los `preset` admitidos. `filename/mime/size/tag`, fuzzy/proximity/boost/intervals no están implementados en este perfil y producen errores explícitos. +Los campos son `text`, `body`, `from`, `chat`, `date`, `kind`, `has`, `topic`, `in`, `preset`, `filename`, `mime`, `size` y `tag`. Los campos de archivo coinciden cuando al menos un adjunto coincide, por el nombre, tamaño y tipo que informó el mensajero; MAX no informa el tipo, así que allí `mime:` no encuentra nada. `tag:` lee las etiquetas locales del dueño en el mensaje, su chat o su remitente. `date:` también acepta `today`, `yesterday` y desplazamientos como `7d`, contados desde el momento de la consulta. Los operadores difusos, de proximidad, de peso y de intervalos se rechazan con un error explícito. ```sh tg messages search 'invoice AND has:file' --source all --context 2 --json max messages search 'chat:"Project team" AND date:[2026-10-01 TO 2026-10-31]' --timezone Europe/Madrid --json +tg messages search 'filename:*.pdf size>1MB date:7d' --json ``` -`text` compara palabras normalizadas; `body`, el texto original completo. Por ejemplo, `body:/.*invoice.*/` expresa una condición de subcadena del cuerpo. La sintaxis regex tiene un subconjunto admitido: flags, lookaround y backreferences de JavaScript no son equivalentes. +`text` compara palabras normalizadas; `body` se refiere al texto original completo, distinguiendo mayúsculas. `body:/.*invoice.*/` expresa una condición de subcadena. Una fecha sin hora es un día de calendario en la zona elegida; un día superior inclusivo incluye el día entero, y los días de cambio de hora no se suponen de 24 horas. -Una fecha ISO sin hora representa un día de calendario en la zona IANA elegida. Un límite superior inclusivo incluye todo el día; uno exclusivo lo excluye. Los límites contemplan el horario de verano, sin suponer días de exactamente 24 horas. +Cuando todas las ramas exigen palabras, los resultados se ordenan por BM25 sobre el índice de palabras; si no, primero los más nuevos. `--newest` siempre ordena por tiempo. Los empates se resuelven con identificadores que incluyen la cuenta, así que las páginas son estables. -Si el índice no está preparado, la búsqueda estricta solicita mantenimiento en vez de devolver coincidencias aproximadas por subcadena. Se limitan bytes del query, profundidad del árbol, tamaño del autómata, expansión del vocabulario, candidatos, bytes leídos, trabajo y tiempo. Agotar un presupuesto es un error, no un resultado vacío completo. +Si el índice de palabras no está listo, una consulta con palabras falla con `index_not_ready`, la parte ya construida y el comando que lo termina; nunca recurre a buscar subcadenas. La longitud de la consulta, la profundidad del árbol, el tamaño del autómata, la expansión del vocabulario (10 000 palabras), los candidatos, los bytes leídos, el trabajo y el tiempo están acotados. Superar un límite es un error, nunca un resultado recortado en silencio. -La respuesta incluye versiones del query/campos/presets, zona horaria, orden, `wordsReady`, cuentas/chats cubiertos y completitud, incluso sin coincidencias. Actualmente `inventoryComplete: false` y `lastSyncedAt: null` no certifican un inventario completo ni una sincronización reciente. JSONL transmite elementos; JSON conserva esta envoltura. +La respuesta JSON incluye las versiones del lenguaje, los campos y los presets, la zona horaria, el orden, `wordsReady` y la cobertura de cuentas y chats, incluso sin resultados. `lastSyncedAt` es la hora más antigua en que un chat del alcance se descargó por última vez con `store fetch`, o `null` si alguno nunca se descargó. `inventoryComplete` indica que cada cuenta del alcance entregó alguna vez su lista completa de chats; es un dato del pasado, no una promesa sobre el presente. -El descubrimiento legacy y JavaScript `--regex` son modos explícitos independientes. Sus correcciones y fallback no forman parte de la semántica estricta de Lucene. +`messages stats` ejecuta la misma consulta y cuenta cada mensaje una vez, por chat, remitente, día de calendario u hora. Una búsqueda o recuento correcto guarda su consulta y opciones, nunca sus resultados, en un historial local de las 1000 ejecuciones más recientes; las búsquedas guardadas son entradas con nombre que se analizan de nuevo en cada ejecución. -### Formas de las palabras +La búsqueda anterior (`--language legacy`) y el modo JavaScript `--regex` siguen siendo explícitos y separados de la semántica estricta. -La búsqueda estricta compara palabras normalizadas completas, así que otra forma de una palabra es otra palabra. `квартира` no encuentra `квартиру` ni `квартиры`, y `piso` no encuentra `pisos`. Para encontrar las formas hoy, usa un prefijo como `квартир*` o `pis*`. Un prefijo también puede encontrar otras palabras que empiezan igual, como `квартирант` o `piscina`. +### Formas de las palabras -La corrección de erratas legacy no añade formas de las palabras. Solo corrige palabras que el archivo no contiene; una palabra que está en el archivo se busca tal cual. +La búsqueda estricta compara palabras normalizadas completas, así que otra forma de una palabra es otra palabra. `квартира` no encuentra `квартиру`, y `piso` no encuentra `pisos`. Un prefijo como `квартир*` o `pis*` captura las formas, y también puede captar otras palabras con el mismo inicio, como `квартирант` o `piscina`. La corrección de erratas del modo legacy no añade formas: solo corrige palabras que el archivo no contiene. -Medimos tres maneras de encontrar formas de las palabras en treebanks de Universal Dependencies, con 102 palabras por idioma ([resultados de la medición](https://github.com/leemour/cli-messaging/blob/bench/stemming-vs-trigrams/bench/stemming/results.md)). *Encontrados* es la parte de los mensajes relevantes que devuelve la búsqueda. *Correctos* es la parte de los resultados devueltos que son relevantes. +Medimos tres formas de encontrar las variantes en treebanks de Universal Dependencies, con 102 palabras por idioma ([resultados de la medición](https://github.com/leemour/cli-messaging/blob/bench/stemming-vs-trigrams/bench/stemming/results.md)). *Encontrados* es la proporción de mensajes relevantes que devuelve la búsqueda; *correctos* es la proporción de resultados devueltos que son relevantes. | Coincidencia | Ruso: encontrados | Ruso: correctos | Español: encontrados | Español: correctos | | --- | --- | --- | --- | --- | | Palabra exacta | 26% | 99% | 39% | 87% | -| Vecinos por trigramas (escritura parecida) | 69% | 52% | 85% | 44% | -| Raíces de Snowball | 88% | 84% | 96% | 69% | +| Trigramas (escritura parecida) | 69% | 52% | 85% | 44% | +| Raíces Snowball | 88% | 84% | 96% | 69% | -**Previsto, aún no publicado:** un índice de raíces separado, con los stemmers oficiales de Snowball, que obtienen la raíz de cada palabra antes de normalizarla. La idea es activarlo por defecto para palabras simples. Un modo exacto conserva el comportamiento actual: una palabra entre comillas o `--exact` para toda la consulta. Mientras no se publique, la búsqueda estricta funciona como se describe arriba. +El almacén ya guarda un índice aparte de raíces, construido con los stemmers oficiales de Snowball a partir de cada palabra antes de normalizarla. Los stemmers son un ajuste de todo el almacén (`searchStemmers.cyrillic` y `searchStemmers.latin`: ruso, español, inglés o ninguno), y `store reindex` aplica un cambio. La búsqueda estricta todavía no usa ese índice y funciona como se describe arriba. -## 3. Reconstruir discusiones con un grafo +## 3. Reconstruir conversaciones con un grafo de mensajes -`conversations build` procesa un chat guardado en orden cronológico. Cada mensaje puede tener enlaces candidatos a mensajes anteriores. El padre elegido lo incorpora a una discusión existente; sin padre empieza otra. +`conversations build` recorre en orden temporal los mensajes guardados de un chat. Cada mensaje puede tener vínculos candidatos con mensajes anteriores. Un padre elegido lo une a esa conversación; sin padre empieza una nueva. -La prioridad actual: +Precedencia: -1. **Respuesta del proveedor:** una respuesta explícita a un mensaje anterior presente en la construcción tiene prioridad. -2. **Respuesta del agente:** una propuesta válida de tu agente puede elegir un padre anterior o comenzar una nueva conversación. -3. **Reglas:** el candidato con mayor peso derivado de menciones o continuación del mismo autor. +1. **La respuesta del mensajero:** una respuesta explícita a un mensaje anterior presente en la construcción gana. +2. **La respuesta del agente:** una respuesta guardada y válida del propio agente del dueño puede elegir un padre anterior o empezar una conversación nueva. +3. **Reglas:** el candidato de mayor confianza entre las menciones (por @username o por nombre) y la continuación del mismo remitente. -Las menciones examinan hasta 50 mensajes anteriores y respetan los IDs de temas nativos cuando están presentes. La continuación del autor busca entre los últimos 10 mensajes y cinco minutos, también en el mismo tema. Los pesos son heurísticos, no probabilidades calibradas. Las respuestas explícitas pueden apuntar más atrás si su padre existe en el historial cargado. +Las reglas miran como mucho 50 mensajes atrás: en un grupo medido, el 97% de los padres de respuestas estaban dentro de 50. La continuación del mismo remitente mira los últimos 10 mensajes y cinco minutos. Ambas respetan los identificadores nativos de hilo. Los pesos son heurísticas, no probabilidades calibradas. Una respuesta explícita puede apuntar más atrás si su padre está en la construcción. -Es un **grafo de relaciones entre mensajes dentro de un chat**, no un grafo de conocimiento global de personas, proyectos y chats. Un tema nativo restringe heurísticas, pero no equivale a una conversación inferida. Padres ausentes, historial incompleto y discusiones intercaladas pueden producir agrupaciones imperfectas. +Es un **grafo de relaciones entre mensajes dentro de un chat**, no un grafo de conocimiento sobre personas, proyectos y chats. Los padres que faltan, el historial incompleto y las discusiones entrelazadas pueden producir agrupaciones imperfectas. -La construcción es explícita: la sincronización normal no reconstruye todos los chats automáticamente. Una nueva construcción publica nuevos miembros e IDs de conversaciones. La anterior sigue disponible hasta que la sustitución esté preparada; un fallo no debe mostrar una versión escrita a medias. Después de reconstruir, vuelve a listar las conversaciones en vez de tratar sus IDs como identidades permanentes. +La construcción es explícita. Una construcción nueva publica nuevas pertenencias e identificadores de conversación; la anterior sigue legible hasta que la nueva está completa, y una construcción fallida nunca muestra medio resultado. Cada construcción registra la versión de las reglas, para que una versión posterior pueda nombrar los chats construidos con reglas antiguas. -Tu agente puede mejorar las relaciones: consultar estado de lotes, obtener mensajes, proponer padres, guardar la respuesta estructurada y reconstruir. El CLI no invoca un LLM ocultamente para enlazarlos. Las respuestas cuyos mensajes o padres han cambiado/se han eliminado quedan obsoletas y se excluyen de construcciones posteriores. +Vinculación con ayuda del agente: `skill show link-conversations` da al agente del dueño sus instrucciones; `conversations batches status` estima los mensajes, los lotes y el texto; `batches next` entrega una ventana de 10–200 mensajes (50 por defecto) con los mensajes anteriores; `conversations links add` guarda la respuesta JSON del agente, toda o nada. La CLI nunca llama a un modelo para vincular mensajes. Una respuesta que se refiere a un mensaje cambiado o eliminado queda obsoleta y no entra en construcciones posteriores. ```sh tg conversations build --chat "Project team" @@ -102,13 +103,15 @@ tg conversations list --chat "Project team" --json max conversations batches status --chat "Project team" --json ``` -## 4. Embeddings de fragmentos de conversaciones +## 4. Fragmentos y vectores + +La construcción corta cada conversación en los límites de los mensajes en fragmentos de unos 1200 caracteres, con los nombres de los remitentes. Un solo mensaje demasiado largo queda como un fragmento; el límite de entrada del modelo puede truncarlo. Una conversación sin texto no produce fragmentos. -La construcción divide cada conversación en fragmentos de texto de aproximadamente 1.200 caracteres por límites de mensajes, incluyendo los nombres de autores. Un mensaje demasiado largo ocupa su propio fragmento; el límite de entrada del modelo puede truncarlo. Una conversación solo de medios sin texto no crea fragmentos de texto. +Cada fragmento conserva su primer y último mensaje y un hash de su texto. Los vectores se guardan por modelo y hash del texto, así que un texto idéntico reutiliza un vector y una ejecución interrumpida continúa con lo que falta. Los mensajes siguen siendo la prueba; los vectores son una representación derivada para buscar. -Cada fragmento conserva referencias al primer/último mensaje y un hash de contenido. Los vectores se guardan por identidad del modelo y hash: el texto idéntico puede reutilizar un vector y una ejecución interrumpida continúa con lo pendiente. Los mensajes originales son la evidencia; el vector es una representación derivada para recuperar candidatos. +Los modelos locales se ejecutan con ONNX Runtime (WebAssembly) en Node o Bun. El predeterminado es `e5-small` (multilingual-e5-small, 8 bits, 384 dimensiones, MIT). `embeddinggemma` (EmbeddingGemma 300M, 4 bits, 768 dimensiones) encuentra más, pero es unas siete veces más lento y requiere aceptar sus términos. Los archivos de los modelos están fijados y se verifican por suma de control, y nunca se comparan vectores de dos modelos. -El modelo local predeterminado es `e5-small`, con 384 dimensiones. `embeddinggemma` es opcional, tiene 768 y exige aceptar sus condiciones. Los archivos están fijados y se comprueban sus checksums. No se mezclan modelos/dimensiones diferentes en una comparación semántica. +Una sesión local usa `min(8, núcleos)` hilos. `--workers` inicia varias sesiones y reparte los hilos entre ellas; `--threads` fija el total. ```sh tg models text download e5-small @@ -116,54 +119,77 @@ tg conversations embed --chat "Project team" tg conversations embed status --chat "Project team" --json ``` -La ejecución local permanece en el equipo. Los proveedores hosted elegidos explícitamente usan la clave configurada del propietario y reciben el texto seleccionado. Antes de enviar pasajes, el comando estima el trabajo pendiente y pide confirmación, salvo que se confirme mediante `--yes`. Una búsqueda hosted también envía el texto del query al proveedor. La búsqueda ordinaria del archivo no requiere ningún proveedor de embeddings externo. +Un proveedor remoto calcula los vectores con la clave del dueño: OpenAI, o cualquier servidor con la forma `/v1/embeddings` de OpenAI mediante `--base-url` junto con `--model` y `--dims`, como Gemini, Jina, Ollama o LM Studio. Eso envía el texto de los fragmentos, y cada búsqueda envía la pregunta. Antes de enviar nada, `embed` indica los fragmentos, el máximo de tokens y el precio máximo, y espera un sí (`--yes` en scripts, `--max-tokens` como tope). -Antes de calcular el vector, se reconstruye el texto actual y se compara con su hash. Los fragmentos cambiados se omiten y necesitan reconstrucción. El nuevo historial también requiere construcción/embedding para entrar en esta vía. La consulta semántica usa hashes de la construcción actual; no revalida todos los mensajes en cada consulta. Tras una edición pueden persistir representaciones anteriores hasta reconstruir y calcular embeddings. Tener vectores no demuestra frescura del archivo. +Antes de calcular, el texto de cada fragmento se reconstruye a partir de los mensajes actuales y se compara con su hash guardado. Un fragmento cambiado se omite hasta la siguiente construcción. -## 5. Recuperación híbrida de conversaciones +### Mediciones + +100 000 mensajes en tres idiomas, de 5 a 40 palabras cada uno, como un solo chat; AMD Ryzen AI 9 HX 470 (24 hilos), 30 GB de RAM, Linux, Node 24 y Bun 1.3; medido el 2 de octubre de 2026 ([medición](https://github.com/leemour/cli-messaging/blob/v0.148.2/bench/embeddings/README.md)). + +| Paso | Entorno | Tiempo | Memoria máxima | +| --- | --- | --- | --- | +| `conversations build` | Node | 3,3 s | 331 MB | +| `conversations embed`, 1 proceso | Node | 1371 s (31 fragmentos/s) | 1,2 GB | +| `conversations embed`, 3 procesos | Node | 1240 s (34 fragmentos/s) | 2,6 GB | +| `conversations embed`, 1 proceso | Bun | 1329 s (32 fragmentos/s) | 3,2 GB | +| `conversations embed`, 3 procesos | Bun | 1279 s (33 fragmentos/s) | 2,7 GB | -`conversations search` acepta una pregunta natural, no una expresión Lucene. Ejecuta dos ramas: +La ejecución produjo 42 417 fragmentos y el almacén creció 292 MB. Tres procesos solo dieron 1,1× en Node y 1,04× en Bun: una sesión ya usa 8 hilos, y tres comparten más o menos los mismos. Una búsqueda, desde que el recorrido de vectores lee cada página por su clave: unos 0,35 s en un proceso que mantiene el modelo abierto (el servidor MCP) y unos 1,4 s para un comando puntual, de los que cargar el modelo es cerca de 1 s. -- **Significado:** convierte el query con el modelo elegido, examina vectores de las construcciones actuales en la cuenta/chat y conserva el mejor fragmento de cada conversación. -- **Palabras:** extrae palabras, las une con OR para buscar mensajes y los relaciona con sus conversaciones construidas. +## 5. Búsqueda híbrida de conversaciones -Combina las listas mediante reciprocal rank fusion: cada una aporta `1 / (60 + rank)`, comenzando en 1. `by` indica `["meaning"]`, `["words"]` o ambos. `summary` contiene metadatos de conversación (IDs, fechas, número de mensajes), no texto generado. `score` es el coseno del mejor fragmento, no la puntuación de fusión ni la probabilidad de una respuesta correcta; las coincidencias solo de palabras tienen `score: null`. +`conversations search` recibe una pregunta en lenguaje natural, no una consulta Lucene. Ejecuta dos ramas: -Los vectores son blobs de SQLite. La implementación los lee por páginas y calcula similitud en JavaScript: no tiene un índice ANN/vectorial ni una base vectorial separada. Las páginas acotan los vectores cargados a la vez, pero el trabajo total crece con los fragmentos del ámbito. +- **Significado:** convertir la pregunta en vector con el modelo elegido, recorrer los vectores de las construcciones actuales del alcance y quedarse con el mejor fragmento de cada conversación. +- **Palabras:** tomar las palabras de la pregunta, unirlas con OR sobre el índice de palabras y asignar los mensajes encontrados a conversaciones construidas. -Un chat vectorizado solo con otro modelo queda fuera de la rama semántica del modelo elegido y se identifica en `embeddedOnlyElsewhere`; sus conversaciones pueden aportar coincidencias léxicas. Historial descargado sin construcción no aparece como conversación. El comando actual codifica el query incluso para resultados solo léxicos: el modelo elegido debe estar instalado/configurado y su ausencia no activa una búsqueda sin modelo. Puede haber candidatos cercanos aunque el archivo no contenga la respuesta: comprueba los originales. +Las dos listas se combinan con reciprocal rank fusion: cada lista suma `1 / (60 + puesto)` a una conversación, con puestos desde uno. Un resultado indica `by: ["meaning"]`, `["words"]` o ambos. Su `score` es la similitud coseno del mejor fragmento, no la puntuación combinada ni la probabilidad de que responda; un resultado encontrado solo por palabras tiene `score: null`. Su `summary` son metadatos (identificadores, fechas, número de mensajes y remitentes), no texto generado. + +Sin un modelo local descargado, la búsqueda responde solo por palabras e indica `"meaning": "unavailable"`; nunca descarga un modelo ni cambia a uno remoto. Un chat con vectores solo de otro modelo queda fuera de la rama de significado y aparece en `embeddedOnlyElsewhere`. Un chat nunca construido no devuelve nada aquí. + +Los vectores son blobs de SQLite que se recorren por páginas y se puntúan en JavaScript. No hay índice de vecinos aproximados ni base de datos vectorial aparte: el trabajo total crece con los fragmentos del alcance. Siempre aparecen vecinos cercanos aunque el archivo no tenga la respuesta, así que verifica los mensajes originales. + +`conversations related` usa como consulta la media de los vectores actuales de una conversación y busca en todos los chats construidos. No se ejecuta ningún modelo, así que no hace falta descargarlo. ```sh tg conversations search "Where did we discuss renting a flat?" --chat "Project team" --json -max conversations search "What changed in the project budget?" --json +max conversations related "Project team" 4521 --json ``` -CLI y MCP usan servicios compartidos. El esquema MCP de conversation search admite query/chat/since/limit, no todas las opciones model/provider del CLI. MCP ofrece list/show/search; construir y vectorizar pasajes son preparaciones explícitas. Un modelo puede codificar la pregunta, pero ninguna vía genera por sí misma una respuesta de IA. +## 6. Frescura -## 6. Contexto y evidencia para un agente +`conversations status` informa, para cada chat construido, un estado —`ready`, `stale`, `partial`, `words-only` o `not-built`— con los mensajes que la construcción no ha visto (nuevos, editados, eliminados), si se usaron reglas antiguas y los fragmentos cuyo vector está al día, obsoleto o ausente. Sin `--chat` también cuenta los grupos nunca construidos. -Después de recuperar candidatos, abre los mensajes o la conversación originales. Distingue una coincidencia exacta, un candidato semántico, una relación inferida por el grafo y una conclusión del agente. Conserva el localizador, chat, fecha y texto relevante junto a cada afirmación. +- **Ediciones.** Un resultado cuyo fragmento cambió después de calcular el vector se marca `stale: true`; su puntuación corresponde al texto anterior hasta la siguiente construcción y cálculo. +- **Eliminaciones.** Un mensaje eliminado conserva su fila, para que una sincronización posterior no lo traiga de vuelta, pero pierde su texto: en la fila, en la copia de búsqueda, en el historial de ediciones, en las transcripciones y en todos los vectores de los fragmentos que lo contenían, salvo que un fragmento actual sin mensajes eliminados use el mismo texto. +- **Puesta al día.** `conversations build` y `embed` sin `--chat`, y `conversations search --refresh`, construyen y calculan en este equipo los chats que cambiaron o nunca se construyeron: por defecto, como mucho 20 chats y 2000 fragmentos por ejecución (`--max-chats`, `--max-chunks`). Nunca descargan un modelo; un modelo remoto solo se usa por chat y con el sí del dueño. MCP ofrece la misma puesta al día como `conversations_refresh`. -El `--context` cronológico no recorre el grafo. `messages links` explica relaciones y la cadena de padres; `conversations show` lee los mensajes agrupados. Los paquetes de evidencia compartidos añaden identidades, fingerprints, límites y metadatos de truncamiento. Una página de evidencia tampoco prueba que el historial esté completo. +## 7. Pruebas y contexto para un agente -Leer/buscar/contexto no marca mensajes como leídos ni envía mensajes. Descargar historial, construir, vectorizar pasajes y usar proveedores externos son acciones explícitas separadas. Un resumen debe declarar datos incompletos/obsoletos en vez de tratar relevancia como prueba. +Después de buscar, abre el mensaje original o la conversación reconstruida. Una respuesta debe distinguir cuatro cosas: una coincidencia exacta de mensaje, un candidato por significado, una relación del grafo y la conclusión del propio agente. Conserva junto a cada afirmación el localizador, el chat, la fecha y el texto relevante alrededor. -## 7. Qué demuestra el playground +`--context` sigue el tiempo, no el grafo. `messages links` explica las relaciones del grafo y la cadena de padres elegida; `conversations show` lee los mensajes agrupados. Los paquetes de pruebas añaden identidades de fuente, huellas del contenido, límites y datos de truncado; aun así, una página de pruebas no demuestra que el historial de un chat esté completo. -El [playground interactivo](./search-playground.mdx) usa el parser compartido con 12 mensajes ficticios en cuatro chats. Demuestra coincidencias estrictas, sugerencias de campos/valores, filtros reversibles, fechas, Matches/All y contexto original. Los errores destacan el segmento incorrecto y conservan los últimos resultados válidos. +Leer, buscar y ver el contexto nunca marca mensajes como leídos ni envía nada. Descargar historial, construir, calcular vectores y usar proveedores remotos son acciones separadas y explícitas. -El navegador no abre cuentas, reconstruye conversaciones, calcula embeddings ni invoca IA real. Sus resúmenes citados son ejemplos preparados, visibles cuando se recupera su evidencia. La interfaz está traducida y los mensajes ficticios siguen en inglés. El motor real admite más campos que el evaluador de ejemplo. +## 8. Qué muestra el playground -## Código y lecturas adicionales +El [playground interactivo](./search-playground.mdx) ejecuta el analizador de consultas compartido sobre 36 mensajes ficticios en seis chats. Muestra la coincidencia estricta, sugerencias de campos y valores, filtros que se pueden quitar, fechas, predicados de archivos y enlaces y el contexto alrededor de un resultado. Una edición inválida conserva los últimos resultados válidos y marca el error de sintaxis. -Fuente examinada: [`cli-messaging` 0.140.0, commit `680d22e`](https://github.com/leemour/cli-messaging/tree/680d22e). Son referencias de implementación, no garantías de rendimiento ni afirmaciones sobre un binary antiguo instalado. +El navegador no abre ninguna cuenta ni ejecuta construcción de conversaciones, vectores ni IA en vivo. Sus resúmenes breves son ejemplos preparados que solo se muestran cuando se recuperan sus pruebas. La interfaz está traducida; los mensajes de ejemplo siguen en inglés. + +## Mapa del código + +Código del motor en [`cli-messaging` v0.148.2](https://github.com/leemour/cli-messaging/tree/v0.148.2). Son referencias de implementación, no promesas de rendimiento. | Responsabilidad | Implementación | | --- | --- | -| Parse, validation, scope, coverage | [Servicio de mensajes](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/messages-search.ts), [ejecutor SQLite Lucene](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/lucene.ts) | -| Grafo y ciclo de construcción | [Reglas](https://github.com/leemour/cli-messaging/blob/680d22e/src/conversations/link.ts), [servicio de conversaciones](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/conversations.ts), [persistencia](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/conversations.ts) | -| Fragmentos y vectores | [División](https://github.com/leemour/cli-messaging/blob/680d22e/src/conversations/chunks.ts), [almacenamiento/similitud](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/vectors.ts) | -| Modelos y fusión | [Servicio de embeddings](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/embeddings.ts), [catálogo](https://github.com/leemour/cli-messaging/blob/680d22e/src/embeddings/models.ts) | -| Contrato de evidencia | [Servicio de evidencia](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/evidence.ts) | - -Consulta la [referencia canónica del lenguaje](https://github.com/leemour/cli-messaging/blob/680d22e/docs/search/query-language.md) para gramática/campos/límites y los archivos de [Telegram](./tg/archive.md) y [MAX](./max/archive.md) para preparación y opciones. +| Análisis, validación, alcance y cobertura | [Servicio de búsqueda de mensajes](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/messages-search.ts) y [ejecutor Lucene para SQLite](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/lucene.ts) | +| Índice de palabras y vocabulario | [Coincidencia de palabras](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/words.ts) y [llenado del índice](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/search-index.ts) | +| Vínculos del grafo y ciclo de construcción | [Reglas de vínculos](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/conversations/link.ts), [servicio de conversaciones](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/conversations.ts) y [persistencia de construcciones](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/conversations.ts) | +| Fragmentos y almacenamiento de vectores | [Fragmentación](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/conversations/chunks.ts) y [almacenamiento y puntuación de vectores](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/vectors.ts) | +| Modelos, combinación y frescura | [Servicio de vectores](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/embeddings.ts), [catálogo de modelos](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/embeddings/models.ts) y [grupo de procesos](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/embeddings/pool.ts) | +| Contrato del paquete de pruebas | [Servicio de pruebas](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/evidence.ts) | + +La gramática exacta, los campos, los presets y los límites están en la [referencia del lenguaje de consulta](https://github.com/leemour/cli-messaging/blob/v0.148.2/docs/search/query-language.md). La preparación y las opciones están en las guías del archivo de [Telegram](./tg/archive.md) y [MAX](./max/archive.md). diff --git a/content/docs/search-architecture.mdx b/content/docs/search-architecture.mdx index 8ae77aa..52fc980 100644 --- a/content/docs/search-architecture.mdx +++ b/content/docs/search-architecture.mdx @@ -1,11 +1,11 @@ --- -title: Search architecture -description: How message search, conversation graphs, embeddings and evidence work together. +title: How search works +description: "The technical side of message search and topic search — indexes, the conversation graph, vectors, ranking, freshness and measurements." --- -WireCat has two retrieval paths over a shared local archive: **strict message search** finds messages satisfying a query, and **conversation search** finds discussions by meaning and shared words. Conversation graphs and embeddings already exist. They are optional processing stages, separate from the Lucene message-search path. +WireCat has two ways to search one local archive. **Message search** finds the messages that satisfy a strict query. **Topic search** finds conversations by meaning and shared words. They share the archive and nothing else: the conversation graph and the vectors are optional stages that message search never reads. -This guide describes the implemented architecture inspected on 4 October 2026. Command details live in the [Telegram search guide](./tg/search.md), [MAX search guide](./max/search.md) and their archive guides. Source links below identify the inspected engine snapshot; they do not mean every installed CLI already has that SDK version. +This is the technical page. For everyday use, read the [Telegram search guide](./tg/search.md) or the [MAX search guide](./max/search.md) and their archive guides. The page describes the engine in [`cli-messaging` 0.148.2](https://github.com/leemour/cli-messaging/tree/v0.148.2), which the development versions of tg and max use. The current releases (tg 0.27, max 0.28) run 0.147.0, which does not have tags, saved searches or the stem settings yet. @@ -13,60 +13,61 @@ This guide describes the implemented architecture inspected on 4 October 2026. C | Question | Entry point | What it returns | Preparation | | --- | --- | --- | --- | -| “Where did Alice say invoice, in October, with a file?” | `messages search` | Matching messages, locators and optional nearby context | Stored history; ready word index for text queries | -| “Where did we discuss renting a flat?” | `conversations search` | Ranked conversations, matching chunks and evidence of why they matched | Conversation build; embeddings for semantic matches | -| “Which discussion does this message belong to?” | `conversations show` / `messages links` | Conversation membership, candidate links and the chosen parent chain | Conversation build | -| “What was around this exact message?” | `messages context` / search `--context` | Chronological neighbors in its account and chat | Stored history; no graph or embedding required | +| “Where did Alice say invoice, in October, with a PDF?” | `messages search` | Matching messages, locators and optional nearby context | Stored history; a ready word index for queries with words | +| “How many messages per day mention the budget?” | `messages stats` | Counts by chat, sender, day or hour | The same as message search | +| “Where did we discuss renting a flat?” | `conversations search` | Ranked conversations, the matching piece and how it matched | A conversation build; embeddings for matches by meaning | +| “What else was about the same thing?” | `conversations related` | Conversations nearest to the one a message is in | Embeddings; no model runs | +| “Which discussion does this message belong to?” | `conversations show` / `messages links` | The conversation and the chosen chain of parents | A conversation build | +| “What was around this exact message?” | `messages context` / search `--context` | Neighbours in time, in the same account and chat | Stored history only | -A native messenger topic, an inferred conversation and a chronological context window are different things. A conversation can join nonadjacent messages; nearby messages may belong to unrelated discussions. +A messenger's native forum topic, an inferred conversation and a window of neighbouring messages are three different things. A conversation can join messages that are far apart; neighbouring messages can belong to unrelated discussions. ## 1. Archive and identity -Reads, explicit history fetches and enabled background capture write messages into a shared SQLite store. Search reads this local copy; it does not silently fetch all remote history. Provider/account/chat/message identifiers keep equally numbered messages in different sources separate. Results expose qualified message locators so an agent can return to the original source. +Reads, explicit history fetches and background capture write messages into one shared SQLite store. Search reads this local copy; it never silently fetches remote history. Provider, account, chat and message identifiers keep equally numbered messages from different sources apart, and every result carries a qualified locator, so an agent can return to the original. -The archive stores original message text and provider metadata separately from normalized search text. It also tracks known history ranges and completeness. Having an index, a built conversation or an embedding does not prove that the entire chat has been downloaded. +The store keeps each message's original text and, separately, its normalized text for search. It also records which ranges of each chat's history it holds. An index, a build or a vector does not prove that a whole chat was downloaded. -Strict message search starts in the active account. `--source all` or `in:all` explicitly widens its scope to accounts held in the store; provider scopes narrow that selection. Names resolve inside the selected scope. Ambiguous authors or chats require a more precise identifier or narrower scope. +Message search starts in the active account. `--source all` or `in:all` widens it to every account in the store; a provider narrows that. Names resolve inside the chosen scope, and an ambiguous author or chat asks for a more precise identifier. Topic search reads one account: its built conversations, optionally one chat and a time boundary. -Conversation search currently searches one account: the active account's built conversations, optionally limited to one chat and a time boundary. It has no equivalent cross-account `--source all` option. Cross-messenger message retrieval does not imply cross-messenger semantic retrieval. +## 2. Strict message search -## 2. Strict Lucene message search +The pipeline: -The execution pipeline is: - -1. Parse the text into a versioned Boolean syntax tree, retaining source positions for errors. The grammar follows a tested Lucene profile; Lucene is the query language, not an Elasticsearch server or a Java Lucene index. -2. Validate fields, values and supported operators against one registry. CLI text and MCP text/AST requests reach the same service. -3. Resolve account scope, cached names and calendar-date boundaries. `topic:` requires a single required chat because topic IDs are chat-local. -4. Compile the tree into parameterized SQLite conditions. Text terms and phrases use the FTS5 word index over normalized text. Metadata predicates constrain the candidate messages. -5. Expand supported text wildcards/regex through the word vocabulary. Full-body patterns and preset detectors run bounded checks over candidates. Regex uses a bounded automaton rather than unrestricted backtracking. -6. Rank or order only satisfying messages, then optionally fetch chronological neighbors in each hit's original account/chat. Relevance ordering never broadens the Boolean result set. +1. Parse the text into a versioned Boolean syntax tree that keeps source positions for errors. The grammar follows a tested profile of Apache Lucene 9.12.3 query syntax. Lucene is the query language here, not a search server or a Java index. +2. Validate fields, values and operators against one registry. CLI text, MCP text and MCP syntax trees reach the same validator and service. +3. Resolve the account scope, stored names and calendar boundaries in the chosen IANA time zone. `topic:` needs exactly one chat, because topic numbers are local to a chat. +4. Compile the tree into parameterized SQLite. Words and phrases use an FTS5 word index over the normalized text; metadata predicates narrow the candidates. +5. Expand wildcards and `text:` regular expressions through the word vocabulary. Full-text patterns and presets run bounded checks over candidates. Regular expressions run as a non-backtracking automaton. +6. Rank or order only messages that match, then optionally fetch neighbours in each hit's own account and chat. Ranking never widens the result. -Fields implemented in the engine are `text`, `body`, `from`, `chat`, `date`, `kind`, `has`, `topic`, `in` and supported `preset` predicates. File-name/MIME/size/tag fields and fuzzy/proximity/boost/interval operators are not implemented in this profile; they return explicit errors. +The fields are `text`, `body`, `from`, `chat`, `date`, `kind`, `has`, `topic`, `in`, `preset`, `filename`, `mime`, `size` and `tag`. File fields match a message when at least one of its attachments matches, by the name, size and type the messenger reported; MAX reports no type, so `mime:` finds nothing there. `tag:` reads the owner's local tags on the message, its chat or its sender. `date:` also takes `today`, `yesterday` and offsets such as `7d`, counted from the moment the query runs. Fuzzy, proximity, boost and interval operators are refused with an explicit error. ```sh tg messages search 'invoice AND has:file' --source all --context 2 --json max messages search 'chat:"Project team" AND date:[2026-10-01 TO 2026-10-31]' --timezone Europe/Madrid --json +tg messages search 'filename:*.pdf size>1MB date:7d' --json ``` -`text` matches normalized words; `body` addresses the original full text. For example, `body:/.*invoice.*/` expresses a full-body substring condition. Regex syntax has its own supported subset; JavaScript flags, lookaround and backreferences are not interchangeable with it. +`text` matches normalized words; `body` addresses the original full text, case-sensitive. `body:/.*invoice.*/` expresses a substring condition. A date without a time is a calendar day in the chosen zone; an inclusive upper day includes the whole day, and days when clocks change are not assumed to last 24 hours. -An ISO date without a time denotes a calendar day in the chosen IANA timezone. An inclusive upper day includes that whole day; an exclusive upper day excludes it. Calendar boundaries account for daylight-saving transitions instead of assuming every day lasts 24 hours. +When every branch requires words, results are ranked by BM25 over the word index; otherwise they come newest first. `--newest` always orders by time. Ties break on account-qualified identifiers, so pages are stable. -If the word index is not ready, strict text search asks for index maintenance rather than returning approximate substring results. Query length, tree depth, automaton size, vocabulary expansion, candidate count, scanned bytes, work and elapsed time are bounded. Exceeding a budget is an error, not a complete empty result. +If the word index is not ready, a query with words fails with `index_not_ready`, the share already built and the command that finishes it; it never falls back to substring matching. Query length, tree depth, automaton size, vocabulary expansion (10,000 words), candidates, scanned bytes, work and time are bounded. Exceeding a budget is an error, never a silently shorter result. -The response includes query/field/preset versions, timezone, ordering, `wordsReady`, account/chat coverage and completeness even when there are no hits. Current coverage includes `inventoryComplete: false` and `lastSyncedAt: null`; do not interpret those as a verified full account inventory or recent synchronization timestamp. JSONL streams items; use JSON when you need this envelope. +The JSON answer includes the language, field and preset versions, the time zone, the order, `wordsReady`, and account and chat coverage, even with no hits. `lastSyncedAt` is the oldest time a chat in scope was last fetched by `store fetch`, or `null` if one never was. `inventoryComplete` says every account in scope has at some point listed all its chats; it is knowledge about the past, not a promise about now. -Legacy discovery and legacy JavaScript `--regex` remain explicit modes. Their correction/fallback behavior is separate from strict Lucene semantics. +`messages stats` runs the same query and counts each matching message once, grouped by chat, sender, calendar day or hour. A successful search or count writes its query and options, never its results, to a local history of the newest 1,000 runs; saved searches are named entries that are parsed again on every run. -### Word forms +Legacy discovery (`--language legacy`) and the legacy JavaScript `--regex` mode remain explicit and separate from strict semantics. -Strict search matches whole normalized words, so another form of a word is another word. `квартира` does not find `квартиру` or `квартиры`, and `piso` does not find `pisos`. To catch the forms today, use a prefix such as `квартир*` or `pis*`. A prefix can also catch other words that start the same way, such as `квартирант` or `piscina`. +### Word forms -The legacy typo correction does not add word forms. It corrects only words that the archive does not contain; a word that is in the archive is searched as it is. +Strict search matches whole normalized words, so another form of a word is another word. `квартира` does not find `квартиру`, and `piso` does not find `pisos`. A prefix such as `квартир*` or `pis*` catches the forms, and can also catch other words that start the same way, such as `квартирант` or `piscina`. Typo correction in legacy mode does not add word forms: it corrects only words the archive does not contain. -We measured three ways of matching word forms on Universal Dependencies treebanks, with 102 words per language ([benchmark results](https://github.com/leemour/cli-messaging/blob/bench/stemming-vs-trigrams/bench/stemming/results.md)). *Found* is the share of relevant messages that the search returns. *Correct* is the share of returned results that are relevant. +We measured three ways of matching word forms on Universal Dependencies treebanks, with 102 words per language ([benchmark results](https://github.com/leemour/cli-messaging/blob/bench/stemming-vs-trigrams/bench/stemming/results.md)). *Found* is the share of relevant messages the search returns; *correct* is the share of returned results that are relevant. | Matching | Russian: found | Russian: correct | Spanish: found | Spanish: correct | | --- | --- | --- | --- | --- | @@ -74,27 +75,27 @@ We measured three ways of matching word forms on Universal Dependencies treebank | Trigram neighbors (similar spelling) | 69% | 52% | 85% | 44% | | Snowball stems | 88% | 84% | 96% | 69% | -**Planned, not shipped yet:** a separate stem index built with the official Snowball stemmers, which take the stem of each word before it is folded into normalized text. It is meant to be on by default for plain words. An exact mode keeps today's behavior: put one word in quotes, or pass `--exact` for the whole query. Until it ships, strict search works as described above. +The store now keeps a separate stem index built with the official Snowball stemmers, taken from each word before it is folded. The stemmers are a store-wide setting (`searchStemmers.cyrillic` and `searchStemmers.latin`: Russian, Spanish, English or none), and `store reindex` applies a change. Strict matching does not use the stem index yet: search behaves as described above. ## 3. Reconstruct discussions with a message graph -`conversations build` processes one chat's stored messages in chronological order. Each message can have candidate links to earlier messages. A chosen parent joins the message to an existing discussion; no parent starts a new discussion. +`conversations build` walks one chat's stored messages in time order. Each message can have candidate links to earlier messages. A chosen parent joins it to that discussion; no parent starts a new one. -The current precedence is: +Precedence: -1. **Provider reply:** an explicit reply to an earlier message present in the build wins. -2. **Agent answer:** a valid stored answer from the user's agent can choose an earlier parent or explicitly start a new conversation. -3. **Rules:** use the highest-confidence candidate from mentions or same-author continuation. +1. **The messenger's reply:** an explicit reply to an earlier message present in the build wins. +2. **The agent's answer:** a valid stored answer from the owner's own agent can choose an earlier parent or start a new conversation. +3. **Rules:** the highest-confidence candidate from mentions (by @username or by name) or from the same sender continuing. -The rules inspect a recent window. Mentions look back up to 50 messages and respect native thread IDs when present. Same-author continuation looks within the last 10 messages and five minutes, also respecting thread IDs. Rule weights are heuristics, not calibrated probabilities of correctness. Explicit replies can point farther back when the parent exists in the loaded chat. +Rules look back at most 50 messages: in a measured group, 97% of reply parents were within 50. Same-sender continuation looks within the last 10 messages and five minutes. Both respect native thread identifiers. Rule weights are heuristics, not calibrated probabilities. An explicit reply can point farther back when its parent is in the build. -This is a **per-chat message relationship graph**, not an entity knowledge graph spanning all people, projects and chats. A native topic helps constrain heuristic links but is not itself an inferred conversation. Missing reply parents, incomplete history and interleaved discussions can produce imperfect grouping. +This is a **per-chat graph of message relationships**, not a knowledge graph of people, projects and chats. Missing reply parents, incomplete history and interleaved discussions can produce imperfect grouping. -Building is explicit; normal archive synchronization does not reconstruct every chat automatically. A new build publishes new conversation memberships and IDs. The previous completed build remains readable until replacement is ready; a failed build must not expose a half-written result. Re-list conversations after rebuilding rather than saving their IDs as permanent identities. +Building is explicit. A new build publishes new memberships and conversation identifiers; the previous build stays readable until the new one is complete, and a failed build never exposes half a result. Each build records its rules version, so a later version can name the chats built with older rules. -Agent-assisted reconstruction is an existing workflow: inspect batch status, obtain a message batch, let your own agent propose parent links, store its structured answer, then rebuild. The CLI does not secretly invoke an LLM to link messages. Answers referring to changed/deleted messages or parents become stale and are excluded from later builds. +Agent-assisted linking: `skill show link-conversations` gives the owner's agent its instructions; `conversations batches status` estimates the messages, batches and text; `batches next` hands out a window of 10–200 messages (50 by default) with the messages before it; `conversations links add` stores the agent's JSON answer, all or nothing. The CLI never calls a model to link messages. An answer that refers to a changed or deleted message becomes stale and is left out of later builds. ```sh tg conversations build --chat "Project team" @@ -102,13 +103,15 @@ tg conversations list --chat "Project team" --json max conversations batches status --chat "Project team" --json ``` -## 4. Embed conversation chunks +## 4. Pieces and vectors + +The build cuts each conversation at message boundaries into pieces of about 1,200 characters, sender names included. A single oversized message stays one piece; the model's input limit may truncate it. A conversation with no text makes no piece. -The build cuts each conversation at message boundaries into text chunks targeting 1,200 characters, including sender names. A single oversized message is kept as one chunk; the embedding model's input limit can truncate it. A media-only conversation with no text creates no text chunk. +Each piece keeps its first and last message and a hash of its text. Vectors are stored by model and text hash, so identical text reuses a vector and an interrupted run resumes with what is left. Messages remain the evidence; vectors are a derived representation for retrieval. -Chunks retain first/last message references and a content hash. Vectors are cached by model identity and content hash, so identical text can reuse a vector and interrupted embedding resumes from the remaining chunks. Original messages remain the evidence; vectors are a derived retrieval representation. +Local models run through ONNX Runtime (WebAssembly) in Node or Bun. The default is `e5-small` (multilingual-e5-small, 8-bit, 384 dimensions, MIT). `embeddinggemma` (EmbeddingGemma 300M, 4-bit, 768 dimensions) finds more but runs about seven times slower and needs its terms accepted. Model files are pinned and checksum-verified, and two models' vectors are never compared with each other. -The default local model is `e5-small`, producing 384-dimensional vectors. The optional `embeddinggemma` model produces 768-dimensional vectors and requires accepting its model terms. Model files are pinned and checksum-verified. Different models/dimensions are never mixed in one semantic comparison. +One local session uses `min(8, cores)` threads. `--workers` starts several sessions with the threads split between them, and `--threads` sets the total. ```sh tg models text download e5-small @@ -116,54 +119,77 @@ tg conversations embed --chat "Project team" tg conversations embed status --chat "Project team" --json ``` -Local embedding runs on the machine. Explicit hosted providers are also available using the owner's configured key; that sends selected text to the provider. The embedding command estimates the remaining work and, for a hosted model, asks for confirmation before embedding passages unless explicitly confirmed with `--yes`. Selecting a hosted provider for a search also sends the query to that provider. Local archive search itself does not require a hosted model. +A remote provider computes vectors with the owner's key: OpenAI, or any server with OpenAI's `/v1/embeddings` shape through `--base-url` with `--model` and `--dims`, such as Gemini, Jina, Ollama or LM Studio. That sends the pieces' text, and each search sends the question. Before sending anything, `embed` states the pieces, the maximum tokens and the maximum price and waits for a yes (`--yes` in scripts, `--max-tokens` as a cap). -Before embedding, chunk text is reconstructed from current messages and checked against its stored hash. Changed chunks are skipped and need a rebuild. New history also requires another build/embedding pass to become available to this path. Existing vectors do not establish archive freshness. Semantic lookup uses current-build chunk hashes; it does not rebuild or revalidate every message at query time. Edits can leave older semantic representations until the next rebuild and embedding pass. +Before embedding, each piece's text is rebuilt from the current messages and checked against its stored hash. A changed piece is skipped until the next build. + +### Measurements + +100,000 messages in three languages, 5–40 words each, as one chat; AMD Ryzen AI 9 HX 470 (24 threads), 30 GB RAM, Linux, Node 24 and Bun 1.3; measured 2 October 2026 ([benchmark](https://github.com/leemour/cli-messaging/blob/v0.148.2/bench/embeddings/README.md)). + +| Step | Runtime | Time | Peak memory | +| --- | --- | --- | --- | +| `conversations build` | Node | 3.3 s | 331 MB | +| `conversations embed`, 1 worker | Node | 1,371 s (31 pieces/s) | 1.2 GB | +| `conversations embed`, 3 workers | Node | 1,240 s (34 pieces/s) | 2.6 GB | +| `conversations embed`, 1 worker | Bun | 1,329 s (32 pieces/s) | 3.2 GB | +| `conversations embed`, 3 workers | Bun | 1,279 s (33 pieces/s) | 2.7 GB | + +The run made 42,417 pieces and grew the store by 292 MB. Three workers gave only 1.1× on Node and 1.04× on Bun: one session already uses 8 threads, and three workers share about the same number. A search, once the vector scan reads each page by its key: about 0.35 s in a process that keeps the model open (the MCP server), about 1.4 s for a one-shot command, of which loading the model is about 1 s. ## 5. Hybrid conversation retrieval -`conversations search` accepts a natural-language question, not a Lucene expression. It runs two branches: +`conversations search` takes a question in natural language, not a Lucene query. It runs two branches: + +- **Meaning:** embed the question with the selected model, scan the vectors of current builds in scope and keep each conversation's best piece. +- **Words:** take the question's words, join them with OR over the word index and map the matching messages to built conversations. -- **Meaning:** embed the question with the selected model, scan vectors belonging to current conversation builds in the selected account/chat, and keep each conversation's best matching chunk. -- **Words:** extract words from the question, join them with OR for message retrieval, and map the matching messages back to built conversations. +The two ranked lists are combined by reciprocal rank fusion: each list adds `1 / (60 + rank)` to a conversation, ranks starting at one. A result reports `by: ["meaning"]`, `["words"]` or both. Its `score` is the best piece's cosine similarity, not the fusion score and not a probability that it answers the question; a result found only by words has `score: null`. Its `summary` is metadata — identifiers, dates, message and sender counts — not generated prose. -The two ranked lists are combined with reciprocal rank fusion: each list contributes `1 / (60 + rank)` to a conversation, with ranks starting at one. A result can report `by: ["meaning"]`, `["words"]` or both. Its `summary` contains conversation metadata (IDs, dates, message count), not generated prose. Its `score` is the best chunk's cosine similarity, not the fusion score or a probability that an answer is correct; a word-only result has `score: null`. +With no local model downloaded, search answers by words alone and reports `"meaning": "unavailable"`; it never downloads a model or switches to a remote one. A chat embedded only with another model is left out of the meaning branch and named in `embeddedOnlyElsewhere`. A chat that was never built returns nothing here. -Vectors live in SQLite blobs. The current implementation scans them in pages and scores them in JavaScript; it does **not** use an approximate-nearest-neighbor/vector index or a separate vector database. Paging bounds the vectors loaded at once, but total scan work grows with the chunks in scope. +Vectors are SQLite blobs scanned in pages and scored in JavaScript. There is no approximate-nearest-neighbour index and no separate vector database: total work grows with the pieces in scope. Nearest neighbours appear even when the archive holds no answer, so verify the original messages. -A chat embedded only with a different model is excluded from this model's semantic branch and named in `embeddedOnlyElsewhere`; it may still supply word matches through built conversations. Downloaded history without a build does not appear as a conversation result. The current command still encodes the query even for word-only results, so its selected model must be installed/configured; a missing model does not automatically switch to a model-free search. Nearest-neighbor candidates can appear even when the archive contains no answer: verify the original messages. +`conversations related` uses the mean of a conversation's current vectors as the query and searches every built chat. No model runs, so it needs no download. ```sh tg conversations search "Where did we discuss renting a flat?" --chat "Project team" --json -max conversations search "What changed in the project budget?" --json +max conversations related "Project team" 4521 --json ``` -CLI and MCP expose conversation retrieval through the same shared services. The shared MCP conversation-search schema accepts query, chat, since and limit; it does not expose every CLI model/provider option. MCP offers conversation list/show/search; building and passage embedding remain explicit preparation steps. An embedding model may run to encode a search query, but neither search path itself writes an AI answer. +## 6. Freshness -## 6. Evidence and context for an agent +`conversations status` reports, for each built chat, a state — `ready`, `stale`, `partial`, `words-only` or `not-built` — with the messages the build has not seen (new, edited, deleted), whether the build used older rules, and the pieces whose vector is current, stale or missing. Without `--chat` it also counts the group chats never built. -After retrieval, open the original message or reconstructed conversation. Distinguish four things in an answer: an exact message match, a semantic candidate, a graph-derived relationship and a conclusion drawn by the agent. Keep the message locator, chat, date and relevant surrounding text with the claim. +- **Edits.** A result whose piece changed after it was embedded is marked `stale: true`; its score is for the old text until the next build and embedding. +- **Deletions.** A deleted message keeps its row, so a later sync cannot bring it back, but loses its text: in the row, the search copy, the edit history, transcripts and every vector of a piece that held it, unless a current piece without a deleted message still uses the same text. +- **Catch-up.** `conversations build` and `embed` without `--chat`, and `conversations search --refresh`, build and embed on this machine the chats that changed or were never built: at most 20 chats and 2,000 pieces a run by default (`--max-chats`, `--max-chunks`). They never download a model; a remote model is only ever used per chat, with the owner's yes. MCP offers the same catch-up as `conversations_refresh`. -Chronological `--context` does not follow a graph. `messages links` explains graph relations and the chosen parent chain; `conversations show` reads the grouped messages. Shared evidence packets additionally preserve source identities, content fingerprints, bounds and truncation metadata for agent workflows. An evidence page still cannot prove full chat history. +## 7. Evidence and context for an agent -Read/search/context operations do not mark messages read or send messages. History fetching, reconstruction, passage embedding and hosted-provider use are explicit separate actions. A summary should disclose incomplete or stale input instead of treating retrieval relevance as proof. +After retrieval, open the original message or the reconstructed conversation. An answer should keep four things apart: an exact message match, a candidate by meaning, a relationship from the graph and the agent's own conclusion. Keep the locator, chat, date and the relevant surrounding text with each claim. -## 7. What the playground demonstrates +`--context` follows time, not the graph. `messages links` explains graph relations and the chosen chain of parents; `conversations show` reads the grouped messages. Evidence packets add source identities, content fingerprints, bounds and truncation metadata; even so, a page of evidence cannot prove a chat's full history. -The [interactive playground](./search-playground.mdx) uses the shared query parser over 12 fictional messages in four chats. It demonstrates strict matching, contextual field/value suggestions, removable filters, dates, Matches/All browsing and original-message context. Invalid edits retain the last valid results while marking the syntax problem. +Reading, searching and context never mark messages read or send anything. Fetching history, building, embedding and remote providers are separate, explicit actions. -The browser does not open an account or run conversation reconstruction, embeddings or live AI. Its short cited summaries are prepared examples shown only when their supporting evidence is retrieved. Its interface is translated; the sample messages remain English. The real engine supports more fields than the sample evaluator. +## 8. What the playground demonstrates -## Source map and further reading +The [interactive playground](./search-playground.mdx) runs the shared query parser over 36 fictional messages in six chats. It shows strict matching, field and value suggestions, removable filters, dates, file and link predicates, and context around a hit. An invalid edit keeps the last valid results and marks the syntax problem. -Inspected source: [`cli-messaging` 0.140.0, commit `680d22e`](https://github.com/leemour/cli-messaging/tree/680d22e). These are implementation references, not performance promises or a claim about an older installed binary. +The browser opens no account and runs no conversation build, embedding or live AI. Its short summaries are prepared examples shown only when their evidence is retrieved. The interface is translated; the sample messages stay in English. + +## Source map + +Engine source at [`cli-messaging` v0.148.2](https://github.com/leemour/cli-messaging/tree/v0.148.2). These are implementation references, not performance promises. | Responsibility | Implementation | | --- | --- | -| Parse, validate, scope and coverage | [Message search service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/messages-search.ts) and [Lucene SQLite executor](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/lucene.ts) | -| Graph links and build lifecycle | [Link rules](https://github.com/leemour/cli-messaging/blob/680d22e/src/conversations/link.ts), [conversation service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/conversations.ts) and [build persistence](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/conversations.ts) | -| Chunk text and vector cache | [Chunking](https://github.com/leemour/cli-messaging/blob/680d22e/src/conversations/chunks.ts) and [vector storage/scoring](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/vectors.ts) | -| Semantic model and rank fusion | [Embedding service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/embeddings.ts) and [model catalog](https://github.com/leemour/cli-messaging/blob/680d22e/src/embeddings/models.ts) | -| Evidence packet contract | [Evidence service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/evidence.ts) | - -For exact grammar, supported fields and limits, use the [canonical query-language reference](https://github.com/leemour/cli-messaging/blob/680d22e/docs/search/query-language.md). For preparation and CLI options, use [Telegram archive](./tg/archive.md) and [MAX archive](./max/archive.md). +| Parse, validate, scope and coverage | [Message search service](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/messages-search.ts) and [Lucene SQLite executor](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/lucene.ts) | +| Word index and vocabulary | [Word matching](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/words.ts) and [index filling](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/search-index.ts) | +| Graph links and build lifecycle | [Link rules](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/conversations/link.ts), [conversation service](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/conversations.ts) and [build persistence](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/conversations.ts) | +| Pieces and vector storage | [Chunking](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/conversations/chunks.ts) and [vector storage and scoring](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/vectors.ts) | +| Models, fusion and freshness | [Embedding service](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/embeddings.ts), [model catalog](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/embeddings/models.ts) and [worker pool](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/embeddings/pool.ts) | +| Evidence packet contract | [Evidence service](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/evidence.ts) | + +For the exact grammar, fields, presets and limits, read the [query-language reference](https://github.com/leemour/cli-messaging/blob/v0.148.2/docs/search/query-language.md). For preparation and options, read the [Telegram archive](./tg/archive.md) and [MAX archive](./max/archive.md) guides. diff --git a/content/docs/search-architecture.ru.mdx b/content/docs/search-architecture.ru.mdx index 4e2bf79..a45859f 100644 --- a/content/docs/search-architecture.ru.mdx +++ b/content/docs/search-architecture.ru.mdx @@ -1,100 +1,101 @@ --- -title: Архитектура поиска -description: Как связаны поиск сообщений, граф разговоров, embeddings и доказательные источники. +title: Как устроен поиск +description: "Техническая сторона поиска сообщений и поиска по темам — индексы, граф разговоров, векторы, порядок результатов, свежесть и замеры." --- -В WireCat есть два пути поиска по общему локальному архиву: **строгий поиск сообщений** находит сообщения, соответствующие запросу, а **поиск разговоров** — обсуждения, близкие по смыслу и словам. Граф связей и embeddings уже реализованы. Это дополнительные этапы обработки, отдельные от поиска сообщений на языке Lucene. +В WireCat два способа искать в одном локальном архиве. **Поиск сообщений** находит сообщения, которые подходят под строгий запрос. **Поиск по темам** находит разговоры по смыслу и общим словам. Общий у них только архив: граф разговоров и векторы — необязательные этапы, которые поиск сообщений не читает. -Описание основано на реализации, прочитанной 4 октября 2026 года. Подробности команд — в руководствах по поиску [Telegram](./tg/search.md) и [MAX](./max/search.md) и страницах архива. Ссылки на исходники фиксируют исследованную версию движка, а не обещают, что она уже установлена в каждом CLI. +Это техническая страница. Для повседневной работы читайте руководства по поиску [Telegram](./tg/search.md) и [MAX](./max/search.md) и страницы архива. Здесь описан движок [`cli-messaging` 0.148.2](https://github.com/leemour/cli-messaging/tree/v0.148.2), на котором работают версии tg и max в разработке. Текущие выпуски (tg 0.27, max 0.28) работают на 0.147.0: там ещё нет меток, сохранённых поисков и настройки основ слов. ## Какой путь выбрать -| Вопрос | Команда | Результат | Подготовка | +| Вопрос | Команда | Что возвращает | Что нужно заранее | | --- | --- | --- | --- | -| «Где Alice упомянула invoice в октябре и приложила файл?» | `messages search` | Совпавшие сообщения, locators и соседний контекст | Сохранённая история; готовый word index для текстовых условий | -| «Где мы обсуждали аренду квартиры?» | `conversations search` | Разговоры по релевантности, совпавшие chunks и способ поиска | Построенные разговоры; embeddings для поиска по смыслу | -| «К какому разговору относится это сообщение?» | `conversations show` / `messages links` | Участники разговора, возможные связи и выбранная цепочка родителей | Построенные разговоры | -| «Что писали рядом с этим сообщением?» | `messages context` / `--context` | Хронологические соседи в том же аккаунте и чате | Сохранённая история; граф и embeddings не нужны | +| «Где Алиса писала про счёт, в октябре, с PDF?» | `messages search` | Подходящие сообщения, их адреса и при желании соседние сообщения | Сохранённая история; готовый индекс слов для запросов со словами | +| «Сколько сообщений в день упоминают бюджет?» | `messages stats` | Числа по чатам, отправителям, дням или часам | То же, что для поиска сообщений | +| «Где мы обсуждали аренду квартиры?» | `conversations search` | Разговоры по порядку, подходящий кусок и как он найден | Построенные разговоры; векторы для поиска по смыслу | +| «Что ещё было о том же?» | `conversations related` | Разговоры, ближайшие к тому, где лежит сообщение | Векторы; модель не запускается | +| «К какому обсуждению относится это сообщение?» | `conversations show` / `messages links` | Разговор и выбранная цепочка родителей | Построенные разговоры | +| «Что было вокруг этого сообщения?» | `messages context` / `--context` у поиска | Соседи по времени в том же аккаунте и чате | Только сохранённая история | -Тема мессенджера, восстановленный разговор и окно соседних сообщений — разные вещи. Разговор может связывать сообщения, не стоящие рядом; хронологические соседи могут обсуждать другое. +Ветка форума в мессенджере, найденный разговор и окно соседних сообщений — три разные вещи. Разговор может соединять далёкие друг от друга сообщения, а соседние сообщения могут относиться к разным обсуждениям. -## 1. Архив и идентичность +## 1. Архив и адреса сообщений -Чтение, явный fetch истории и включённый сбор новых сообщений сохраняют данные в общей SQLite. Поиск читает эту копию и не скачивает незаметно всю удалённую историю. Provider/account/chat/message позволяют отличить одинаковые номера сообщений в разных источниках. В ответах есть квалифицированные locators для перехода к оригиналу. +Чтение, явная загрузка истории и фоновая запись кладут сообщения в одно общее хранилище SQLite. Поиск читает эту локальную копию и никогда сам не догружает историю из сети. Провайдер, аккаунт, чат и сообщение разделяют сообщения с одинаковыми номерами из разных источников, а у каждого результата есть полный адрес, по которому агент вернётся к оригиналу. -Исходный текст и metadata провайдера хранятся отдельно от нормализованного текста индекса. Архив учитывает известные диапазоны истории и полноту. Готовый индекс, разговор или вектор не доказывает, что скачан весь чат. +Хранилище держит исходный текст каждого сообщения и отдельно — нормализованный текст для поиска. Ещё оно помнит, какие участки истории каждого чата у него есть. Индекс, построение или вектор не доказывают, что чат скачан целиком. -Строгий поиск начинает с активного аккаунта. `--source all` или `in:all` явно расширяет область до аккаунтов, сохранённых в store; provider scope сужает её. Имена разрешаются внутри выбранной области. Неоднозначные авторы и чаты требуют точного идентификатора или более узкого scope. +Поиск сообщений начинается с активного аккаунта. `--source all` или `in:all` расширяет его до всех аккаунтов хранилища, провайдер сужает выбор. Имена ищутся внутри выбранного охвата, а неоднозначный автор или чат требует более точного идентификатора. Поиск по темам читает один аккаунт: его построенные разговоры, при желании один чат и границу по времени. -Поиск разговоров сейчас ограничен одним аккаунтом: построенными разговорами активного аккаунта, при необходимости одним чатом и временем. Аналога `--source all` у него нет. Межмессенджерный поиск сообщений не означает межмессенджерный semantic search. +## 2. Строгий поиск сообщений -## 2. Строгий поиск сообщений на языке Lucene +Порядок работы: -Путь выполнения: - -1. Текст разбирается в versioned Boolean AST с позициями для ошибок. Используется проверенный профиль грамматики Lucene; это язык запросов, а не сервер Elasticsearch или Java-индекс Lucene. -2. Один registry проверяет поля, значения и операторы. CLI text и MCP text/AST обращаются к одному service. -3. Разрешаются accounts, локальные имена и календарные даты. `topic:` требует одного обязательного чата: номера тем локальны внутри чата. -4. AST превращается в параметризованные SQLite predicates. Термы и фразы используют FTS5 word index по нормализованному тексту; metadata ограничивает сообщения-кандидаты. -5. Wildcard/regex по словам раскрываются через словарь индекса. Проверки полного текста и preset detectors выполняются с ограниченными бюджетами. Regex использует ограниченный автомат, а не произвольный backtracking. -6. Ранжируются только удовлетворяющие запросу сообщения; затем при необходимости читаются их хронологические соседи в исходном аккаунте/чате. Ранжирование не расширяет Boolean множество. +1. Разобрать текст в версионированное логическое дерево, которое помнит позиции для сообщений об ошибках. Грамматика следует проверенному профилю синтаксиса запросов Apache Lucene 9.12.3. Lucene здесь — язык запросов, а не поисковый сервер и не индекс Java. +2. Проверить поля, значения и операторы по одному реестру. Текст из CLI, текст и деревья из MCP проходят один и тот же проверяющий код и сервис. +3. Определить охват аккаунтов, сохранённые имена и календарные границы в выбранном поясе IANA. `topic:` требует ровно одного чата: номера веток действуют внутри чата. +4. Собрать из дерева параметризованный SQL для SQLite. Слова и фразы ищутся по индексу слов FTS5 над нормализованным текстом, условия метаданных сужают кандидатов. +5. Раскрыть шаблоны и регулярные выражения `text:` через словарь слов. Шаблоны по полному тексту и preset проверяют кандидатов в пределах бюджета. Регулярные выражения работают как автомат без отката. +6. Упорядочить только подходящие сообщения и при желании добавить соседей в аккаунте и чате каждого найденного. Порядок никогда не расширяет результат. -Реализованы `text`, `body`, `from`, `chat`, `date`, `kind`, `has`, `topic`, `in` и поддержанный набор `preset`. `filename/mime/size/tag`, fuzzy/proximity/boost/intervals в этом профиле не реализованы и дают явную ошибку. +Поля: `text`, `body`, `from`, `chat`, `date`, `kind`, `has`, `topic`, `in`, `preset`, `filename`, `mime`, `size` и `tag`. Поля файлов подходят сообщению, если подходит хотя бы одно вложение — по имени, размеру и типу, которые сообщил мессенджер; MAX тип не сообщает, поэтому там `mime:` ничего не находит. `tag:` читает локальные метки владельца на сообщении, его чате или отправителе. `date:` понимает и `today`, `yesterday`, и сдвиги вроде `7d` — от момента запроса. Нечёткий поиск, близость слов, веса и интервалы отвергаются явной ошибкой. ```sh tg messages search 'invoice AND has:file' --source all --context 2 --json max messages search 'chat:"Project team" AND date:[2026-10-01 TO 2026-10-31]' --timezone Europe/Madrid --json +tg messages search 'filename:*.pdf size>1MB date:7d' --json ``` -`text` сопоставляет нормализованные слова, `body` — весь исходный текст. `body:/.*invoice.*/` задаёт поиск вхождения в полном тексте. Regex имеет свой поддержанный subset; JavaScript flags, lookaround и backreferences не являются его аналогами. +`text` сравнивает нормализованные слова, `body` — исходный текст целиком, с учётом регистра. `body:/.*invoice.*/` задаёт условие на подстроку. Дата без времени — календарный день в выбранном поясе; включающая верхняя граница включает день целиком, и день перехода на летнее время не считается равным 24 часам. -ISO date без времени означает календарный день в выбранном IANA timezone. Inclusive upper day включает весь день, exclusive исключает его. Границы учитывают DST, а не предполагают ровно 24 часа в каждом дне. +Когда все ветви запроса требуют слов, результаты упорядочены по BM25 над индексом слов; иначе — сначала новые. `--newest` всегда упорядочивает по времени. Равенство решают полные идентификаторы с аккаунтом, поэтому страницы стабильны. -Неготовый word index требует обслуживания: строгий поиск не заменяет его подстрочной выдачей. Ограничены размер запроса, глубина AST, автомат, раскрытие словаря, кандидаты, прочитанные байты, работа и время. Исчерпание бюджета — ошибка, а не полный пустой результат. +Если индекс слов не готов, запрос со словами завершается ошибкой `index_not_ready` с долей готового и командой, которая его достроит; на поиск подстрок он не переходит. Длина запроса, глубина дерева, размер автомата, раскрытие словаря (10 000 слов), число кандидатов, прочитанные байты, работа и время ограничены. Превышение — ошибка, а не молча укороченный результат. -Ответ содержит версии query/fields/presets, timezone, порядок, `wordsReady`, accounts/chat coverage и completeness даже при нуле совпадений. Сейчас `inventoryComplete: false`, `lastSyncedAt: null`: это не подтверждение полного списка чатов или недавней синхронизации. JSONL передаёт items; для полной оболочки используйте JSON. +Ответ в JSON содержит версии языка, полей и preset, часовой пояс, порядок, `wordsReady` и охват аккаунтов и чатов — даже без совпадений. `lastSyncedAt` — самое старое время последней загрузки чата из охвата через `store fetch`, или `null`, если хотя бы один не загружался. `inventoryComplete` говорит, что каждый аккаунт охвата когда-то передал полный список чатов; это знание о прошлом, а не обещание о настоящем. -Legacy discovery и JavaScript `--regex` — отдельные явно выбранные режимы. Их исправление опечаток и fallback не входят в строгую семантику Lucene. +`messages stats` выполняет тот же запрос и считает каждое подходящее сообщение один раз — по чатам, отправителям, календарным дням или часам. Успешный поиск или подсчёт записывает запрос и опции, но не результаты, в локальную историю последних 1000 запусков; сохранённые поиски — именованные записи, которые разбираются заново при каждом запуске. -### Формы слов +Прежний поиск (`--language legacy`) и прежний режим JavaScript `--regex` остаются явными и отдельными от строгого поиска. -Строгий поиск сравнивает целые нормализованные слова, поэтому другая форма слова для него — другое слово. `квартира` не находит `квартиру` и `квартиры`, а `piso` не находит `pisos`. Чтобы найти формы сейчас, используйте префикс: `квартир*` или `pis*`. Префикс может найти и другие слова с тем же началом, например `квартирант` или `piscina`. +### Формы слов -Legacy-исправление опечаток не добавляет формы слов. Оно исправляет только слова, которых нет в архиве; слово, которое в архиве есть, ищется как есть. +Строгий поиск сравнивает целые нормализованные слова, поэтому другая форма слова — другое слово. `квартира` не найдёт «квартиру», а `piso` — «pisos». Начало слова, `квартир*` или `pis*`, ловит формы, но может поймать и другие слова с тем же началом, например «квартирант» или «piscina». Исправление опечаток в режиме legacy форм не добавляет: оно исправляет только слова, которых в архиве нет. -Мы измерили три способа сопоставлять формы слов на трибанках Universal Dependencies, по 102 слова на язык ([результаты замера](https://github.com/leemour/cli-messaging/blob/bench/stemming-vs-trigrams/bench/stemming/results.md)). *Найдено* — доля релевантных сообщений, которые возвращает поиск. *Верно* — доля результатов, которые действительно релевантны. +Мы сравнили три способа находить формы слов на корпусах Universal Dependencies, по 102 слова на язык ([результаты замеров](https://github.com/leemour/cli-messaging/blob/bench/stemming-vs-trigrams/bench/stemming/results.md)). *Найдено* — доля нужных сообщений, которые вернул поиск; *верно* — доля нужных среди возвращённых. | Сопоставление | Русский: найдено | Русский: верно | Испанский: найдено | Испанский: верно | | --- | --- | --- | --- | --- | | Точное слово | 26% | 99% | 39% | 87% | -| Соседи по триграммам (похожее написание) | 69% | 52% | 85% | 44% | +| Триграммы (похожее написание) | 69% | 52% | 85% | 44% | | Основы Snowball | 88% | 84% | 96% | 69% | -**Запланировано, ещё не выпущено:** отдельный индекс основ слов на официальных стеммерах Snowball; основа берётся до того, как слово приводится к нормализованному виду. Для обычных слов его планируется включить по умолчанию. Точный режим сохраняет нынешнее поведение: слово в кавычках или `--exact` для всего запроса. Пока индекс не выпущен, строгий поиск работает так, как описано выше. +Хранилище теперь держит отдельный индекс основ слов, построенный официальными стеммерами Snowball по слову до нормализации. Стеммеры — настройка всего хранилища (`searchStemmers.cyrillic` и `searchStemmers.latin`: русский, испанский, английский или выключено), а применяет изменение `store reindex`. Строгий поиск этот индекс пока не использует и работает, как описано выше. -## 3. Граф связей и восстановление разговоров +## 3. Разговоры как граф сообщений -`conversations build` читает один сохранённый чат в хронологическом порядке. Сообщение может иметь несколько candidate links к более ранним сообщениям. Выбранный родитель присоединяет его к существующему разговору; отсутствие родителя начинает новый. +`conversations build` проходит сохранённые сообщения одного чата по времени. У каждого сообщения могут быть кандидаты-связи с более ранними. Выбранный родитель присоединяет сообщение к его обсуждению; без родителя начинается новое. -Приоритет сейчас такой: +Порядок: -1. **Reply провайдера:** явный ответ на более раннее сообщение, имеющееся в build, побеждает. -2. **Ответ агента:** сохранённый корректный ответ вашего агента может выбрать предыдущего родителя или начать новый разговор. -3. **Правила:** наиболее уверенный кандидат по упоминанию или продолжению того же автора. +1. **Ответ в мессенджере:** явный ответ на более раннее сообщение, которое есть в построении, побеждает. +2. **Ответ агента:** действующий сохранённый ответ собственного агента владельца может выбрать более раннего родителя или начать новый разговор. +3. **Правила:** самый уверенный кандидат из упоминаний (по @username или по имени) или из продолжения тем же отправителем. -Упоминания смотрят назад до 50 сообщений и учитывают native thread ID, если он сохранён. Продолжение автора ищется в последних 10 сообщениях и пяти минутах, также в том же thread. Веса правил — эвристики, а не калиброванная вероятность правильности. Явный reply может уходить дальше, если родитель есть в загруженном чате. +Правила смотрят назад не дальше 50 сообщений: в измеренной группе 97% родителей ответов были в пределах 50. Продолжение тем же отправителем ищется в последних 10 сообщениях и пяти минутах. Оба правила учитывают родные идентификаторы веток. Веса правил — эвристика, а не откалиброванная вероятность. Явный ответ может указывать дальше, если его родитель есть в построении. -Это **граф связей сообщений внутри одного чата**, а не knowledge graph всех людей, проектов и чатов. Native topic ограничивает эвристики, но не равен восстановленному разговору. Пропущенные reply parents, неполная история и перемешанные обсуждения могут ухудшать группировку. +Это **граф связей сообщений внутри чата**, а не граф знаний о людях, проектах и чатах. Отсутствующие родители ответов, неполная история и переплетённые обсуждения дают неидеальное разбиение. -Build запускается явно; обычная синхронизация не перестраивает автоматически все чаты. Новый build публикует новые memberships и conversation IDs. Предыдущий законченный build доступен до готовности замены; ошибка не должна показывать половину нового результата. После rebuild заново получите список разговоров, а не используйте их IDs как вечные идентификаторы. +Построение запускается явно. Новое построение публикует новые составы и номера разговоров; прежнее остаётся читаемым, пока новое не готово целиком, а неудачное построение не показывает половину результата. Каждое построение помнит версию правил, поэтому следующая версия может назвать чаты, построенные по старым. -Агент может улучшить связи: посмотреть batch status, получить пачку сообщений, предложить родителей, записать структурированный ответ и запустить rebuild. CLI сам не вызывает LLM для связывания. Ответы, относящиеся к изменённым/удалённым сообщениям или родителям, становятся stale и исключаются из последующих builds. +Связывание с помощью агента: `skill show link-conversations` даёт агенту владельца инструкцию; `conversations batches status` оценивает число сообщений, пачек и объём текста; `batches next` выдаёт окно из 10–200 сообщений (по умолчанию 50) вместе с предыдущими; `conversations links add` сохраняет ответ агента в JSON — целиком или никак. Сам CLI модель для связывания не вызывает. Ответ, который ссылается на изменённое или удалённое сообщение, устаревает и в следующие построения не попадает. ```sh tg conversations build --chat "Project team" @@ -102,13 +103,15 @@ tg conversations list --chat "Project team" --json max conversations batches status --chat "Project team" --json ``` -## 4. Embeddings частей разговора +## 4. Куски и векторы + +Построение режет каждый разговор по границам сообщений на куски примерно по 1200 символов, вместе с именами отправителей. Одно слишком длинное сообщение остаётся одним куском; предел входа модели может его обрезать. Разговор без текста кусков не даёт. -Build разрезает разговор по границам сообщений на text chunks с целевым размером 1 200 символов, включая имя автора. Одно слишком длинное сообщение остаётся отдельным chunk; input limit модели может обрезать его. Разговор только из media без текста не создаёт text chunk. +Каждый кусок помнит первое и последнее сообщение и хеш своего текста. Векторы хранятся по модели и хешу текста, поэтому одинаковый текст переиспользует вектор, а прерванный запуск продолжает с оставшегося. Доказательством остаются сообщения; векторы — производное представление для поиска. -Chunk хранит ссылки на первое/последнее сообщение и hash содержимого. Векторы кешируются по model identity и hash: одинаковый текст может использовать готовый вектор, а прерванный embed продолжает оставшиеся chunks. Оригинальные сообщения остаются доказательствами; вектор — производное представление для поиска. +Локальные модели работают через ONNX Runtime (WebAssembly) в Node или Bun. По умолчанию — `e5-small` (multilingual-e5-small, 8 бит, 384 измерения, MIT). `embeddinggemma` (EmbeddingGemma 300M, 4 бита, 768 измерений) находит больше, но работает примерно в семь раз медленнее и требует принять её условия. Файлы моделей закреплены и проверяются по контрольной сумме, а векторы двух моделей никогда не сравниваются между собой. -Локальная модель по умолчанию — `e5-small`, 384 измерения. Дополнительная `embeddinggemma` — 768 измерений, с обязательным принятием model terms. Файлы моделей зафиксированы и проверяются checksum. Разные модели/размерности в одном semantic comparison не смешиваются. +Одна локальная сессия занимает `min(8, ядра)` потоков. `--workers` запускает несколько сессий и делит потоки между ними, `--threads` задаёт общее число. ```sh tg models text download e5-small @@ -116,54 +119,77 @@ tg conversations embed --chat "Project team" tg conversations embed status --chat "Project team" --json ``` -Локальный embed работает на компьютере. Явно выбранный hosted provider использует настроенный ключ владельца и получает выбранный текст. Для passages команда оценивает оставшуюся работу и запрашивает подтверждение перед внешним embedding, если не указан `--yes`. Hosted search также отправляет provider текст запроса. Обычному поиску архива внешний embedding provider не нужен. +Внешний провайдер считает векторы с ключом владельца: OpenAI или любой сервер с API `/v1/embeddings` в формате OpenAI — через `--base-url` вместе с `--model` и `--dims`, например Gemini, Jina, Ollama или LM Studio. Тогда ему уходит текст кусков, а каждый поиск отправляет вопрос. Прежде чем что-то отправить, `embed` называет число кусков, наибольшее число токенов и наибольшую цену и ждёт «да» (`--yes` в скриптах, `--max-tokens` — предел). -Перед embed chunk text восстанавливается из текущих сообщений и сверяется с hash. Изменённые chunks пропускаются: нужен rebuild. Новая история также требует нового build/embed, чтобы попасть в этот путь. Semantic lookup использует hash текущего build и не перечитывает fingerprints всех сообщений при каждом поиске: после edits до rebuild могут оставаться старые представления. Наличие векторов не доказывает свежесть архива. +Перед подсчётом текст каждого куска собирается заново из текущих сообщений и сверяется с сохранённым хешем. Изменившийся кусок пропускается до следующего построения. + +### Замеры + +100 000 сообщений на трёх языках, по 5–40 слов, одним чатом; AMD Ryzen AI 9 HX 470 (24 потока), 30 ГБ памяти, Linux, Node 24 и Bun 1.3; замерено 2 октября 2026 года ([замеры](https://github.com/leemour/cli-messaging/blob/v0.148.2/bench/embeddings/README.md)). + +| Шаг | Среда | Время | Пик памяти | +| --- | --- | --- | --- | +| `conversations build` | Node | 3,3 с | 331 МБ | +| `conversations embed`, 1 процесс | Node | 1371 с (31 кусок/с) | 1,2 ГБ | +| `conversations embed`, 3 процесса | Node | 1240 с (34 куска/с) | 2,6 ГБ | +| `conversations embed`, 1 процесс | Bun | 1329 с (32 куска/с) | 3,2 ГБ | +| `conversations embed`, 3 процесса | Bun | 1279 с (33 куска/с) | 2,7 ГБ | + +Получилось 42 417 кусков, хранилище выросло на 292 МБ. Три процесса дали всего 1,1× на Node и 1,04× на Bun: одна сессия уже занимает 8 потоков, а трём достаётся примерно столько же. Поиск — после того как просмотр векторов стал читать каждую страницу по ключу — около 0,35 с в процессе, который держит модель открытой (сервер MCP), и около 1,4 с для разовой команды, из них около 1 с уходит на загрузку модели. ## 5. Гибридный поиск разговоров -`conversations search` принимает вопрос обычным языком, не Lucene expression. Две ветки: +`conversations search` принимает вопрос обычными словами, а не запрос Lucene. Он идёт двумя ветками: + +- **Смысл:** превратить вопрос в вектор выбранной моделью, просмотреть векторы текущих построений в охвате и оставить для каждого разговора лучший кусок. +- **Слова:** взять слова вопроса, соединить их через OR по индексу слов и сопоставить найденные сообщения с построенными разговорами. -- **Смысл:** запрос векторизуется выбранной моделью; читаются vectors текущих conversation builds в выбранном account/chat, для разговора сохраняется лучший chunk. -- **Слова:** из вопроса извлекаются слова, объединяются через OR для message search; найденные сообщения переводятся в построенные разговоры. +Два списка объединяет reciprocal rank fusion: каждый список добавляет разговору `1 / (60 + место)`, места считаются с единицы. Результат сообщает `by: ["meaning"]`, `["words"]` или оба. Его `score` — косинусная близость лучшего куска, а не итог объединения и не вероятность, что он отвечает на вопрос; у найденного только по словам `score: null`. Его `summary` — метаданные (номера, даты, число сообщений и отправителей), а не сгенерированный текст. -Списки соединяются reciprocal rank fusion: каждый даёт `1 / (60 + rank)`, начиная с rank 1. Поле `by` равно `["meaning"]`, `["words"]` или обоим. `summary` — metadata разговора (IDs, даты, число сообщений), не сгенерированный текст. `score` — cosine лучшего chunk, не fusion score и не вероятность верности ответа; у word-only результата `score: null`. +Без скачанной локальной модели поиск отвечает по словам и сообщает `"meaning": "unavailable"`; модель он не скачивает и на внешнюю не переключается. Чат, векторы которого посчитаны только другой моделью, в ветку смысла не попадает и назван в `embeddedOnlyElsewhere`. Чат, который ни разу не построен, здесь ничего не даёт. -Векторы хранятся как SQLite blobs. Сейчас они читаются страницами и сравниваются в JavaScript; ANN/vector index и отдельной vector DB нет. Paging ограничивает одновременно загруженные vectors, но общая работа растёт с числом chunks в scope. +Векторы лежат в SQLite как двоичные значения, просматриваются страницами и оцениваются в JavaScript. Индекса приближённых ближайших соседей и отдельной векторной базы нет: работа растёт с числом кусков в охвате. Ближайшие соседи находятся, даже когда ответа в архиве нет, поэтому проверяйте исходные сообщения. -Чат с embeddings только другой модели исключается из semantic branch этой модели и перечисляется в `embeddedOnlyElsewhere`; его построенные разговоры могут дать word matches. История без build не становится conversation result. Текущая команда векторизует query даже для word-only выдачи: выбранная модель должна быть установлена/настроена, её отсутствие не включает автоматически поиск без модели. Ближайший кандидат может существовать и при отсутствии ответа в архиве: проверьте оригиналы. +`conversations related` берёт как запрос среднее текущих векторов разговора и ищет во всех построенных чатах. Модель не запускается, скачивать её не нужно. ```sh tg conversations search "Where did we discuss renting a flat?" --chat "Project team" --json -max conversations search "What changed in the project budget?" --json +max conversations related "Project team" 4521 --json ``` -CLI и MCP используют общие services. Shared MCP schema conversation search принимает query/chat/since/limit, а не все CLI model/provider options. MCP предлагает list/show/search; build и passage embed остаются явной подготовкой. Для векторизации запроса может работать embedding model, но сами поисковые пути не пишут AI-ответ. +## 6. Свежесть -## 6. Контекст и доказательства для агента +`conversations status` показывает для каждого построенного чата состояние — `ready`, `stale`, `partial`, `words-only` или `not-built` — с сообщениями, которых построение не видело (новые, изменённые, удалённые), признаком старых правил и кусками с актуальным, устаревшим или отсутствующим вектором. Без `--chat` он ещё считает групповые чаты, которые ни разу не построены. -После поиска откройте исходное сообщение или разговор. Различайте точное совпадение, semantic candidate, выведенную графом связь и вывод самого агента. Сохраняйте рядом с утверждением locator, чат, дату и релевантный текст. +- **Правки.** Результат, кусок которого изменился после подсчёта вектора, помечен `stale: true`; его оценка — для старого текста до следующего построения и подсчёта. +- **Удаления.** Удалённое сообщение сохраняет строку, чтобы следующая синхронизация его не вернула, но теряет текст: в строке, в копии для поиска, в истории правок, в расшифровках и во всех векторах кусков, где оно было, — если только тот же текст не использует актуальный кусок без удалённых сообщений. +- **Догонка.** `conversations build` и `embed` без `--chat` и `conversations search --refresh` строят и считают на этом компьютере изменившиеся и ни разу не построенные чаты: по умолчанию не больше 20 чатов и 2000 кусков за запуск (`--max-chats`, `--max-chunks`). Модель они не скачивают; внешняя модель используется только для отдельного чата и с «да» владельца. В MCP та же догонка — `conversations_refresh`. -Хронологический `--context` не ходит по графу. `messages links` объясняет отношения и выбранную parent chain, `conversations show` читает сгруппированные сообщения. Shared evidence packets дополнительно сохраняют source identities, fingerprints, лимиты и truncation metadata. Одна evidence page не доказывает полноту чата. +## 7. Доказательства и контекст для агента -Read/search/context не отмечают сообщения прочитанными и ничего не отправляют. Fetch истории, build, passage embed и hosted provider — отдельные явные действия. Summary должна сообщать неполноту/устаревание входных данных, а не считать релевантность доказательством. +После поиска откройте исходное сообщение или разговор. В ответе стоит различать четыре вещи: точное совпадение сообщения, кандидата по смыслу, связь из графа и собственный вывод агента. Держите рядом с каждым утверждением адрес сообщения, чат, дату и нужный окружающий текст. -## 7. Что показывает playground +`--context` идёт по времени, а не по графу. `messages links` объясняет связи графа и выбранную цепочку родителей; `conversations show` читает сгруппированные сообщения. Пакеты доказательств добавляют идентификаторы источников, отпечатки содержимого, границы и сведения об усечении; но и страница доказательств не доказывает полноту истории чата. -[Интерактивный пример](./search-playground.mdx) использует общий parser на 12 вымышленных сообщениях в четырёх чатах. Он показывает строгие совпадения, подсказки field/value, снимаемые фильтры, даты, Matches/All и исходный context. Ошибка редактирования подсвечивается, предыдущая корректная выдача остаётся. +Чтение, поиск и контекст никогда не помечают сообщения прочитанными и ничего не отправляют. Загрузка истории, построение, подсчёт векторов и внешние провайдеры — отдельные явные действия. -Браузер не открывает аккаунт, не строит разговоры, не вычисляет embeddings и не вызывает живой AI. Короткие cited summaries подготовлены и показываются только при найденных подтверждающих сообщениях. Интерфейс переведён, sample messages — на английском. Реальный движок поддерживает больше полей, чем sample evaluator. +## 8. Что показывает интерактивный пример -## Исходники и дальнейшее чтение +[Интерактивный пример](./search-playground.mdx) запускает общий разборщик запросов на 36 вымышленных сообщениях в шести чатах. Он показывает строгое сопоставление, подсказки полей и значений, снимаемые фильтры, даты, условия на файлы и ссылки и контекст вокруг найденного. Неверная правка сохраняет последние верные результаты и отмечает ошибку синтаксиса. -Исследован [`cli-messaging` 0.140.0, commit `680d22e`](https://github.com/leemour/cli-messaging/tree/680d22e). Ссылки описывают реализацию, не обещают скорость или её наличие в старом установленном binary. +Браузер не открывает аккаунт и не запускает построение разговоров, векторы или живой ИИ. Короткие сводки — заготовленные примеры, которые показываются, только когда найдены их доказательства. Интерфейс переведён, примеры сообщений остаются на английском. -| Область | Реализация | -| --- | --- | -| Parse, validation, scope, coverage | [Message search service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/messages-search.ts), [Lucene SQLite executor](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/lucene.ts) | -| Граф и lifecycle build | [Link rules](https://github.com/leemour/cli-messaging/blob/680d22e/src/conversations/link.ts), [conversation service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/conversations.ts), [build persistence](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/conversations.ts) | -| Chunks и vectors | [Chunking](https://github.com/leemour/cli-messaging/blob/680d22e/src/conversations/chunks.ts), [vector storage/scoring](https://github.com/leemour/cli-messaging/blob/680d22e/src/store/sqlite/vectors.ts) | -| Модели и fusion | [Embedding service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/embeddings.ts), [model catalog](https://github.com/leemour/cli-messaging/blob/680d22e/src/embeddings/models.ts) | -| Evidence contract | [Evidence service](https://github.com/leemour/cli-messaging/blob/680d22e/src/services/evidence.ts) | +## Карта исходников + +Исходники движка — [`cli-messaging` v0.148.2](https://github.com/leemour/cli-messaging/tree/v0.148.2). Это ссылки на реализацию, а не обещания производительности. -Грамматика, поля и пределы — в [канонической справке языка](https://github.com/leemour/cli-messaging/blob/680d22e/docs/search/query-language.md). Подготовка и CLI options — в страницах [архива Telegram](./tg/archive.md) и [архива MAX](./max/archive.md). +| Задача | Реализация | +| --- | --- | +| Разбор, проверка, охват и полнота | [Сервис поиска сообщений](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/messages-search.ts) и [исполнитель Lucene для SQLite](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/lucene.ts) | +| Индекс слов и словарь | [Сопоставление слов](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/words.ts) и [заполнение индекса](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/search-index.ts) | +| Связи графа и жизненный цикл построения | [Правила связей](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/conversations/link.ts), [сервис разговоров](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/conversations.ts) и [хранение построений](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/conversations.ts) | +| Куски и хранение векторов | [Нарезка](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/conversations/chunks.ts) и [хранение и оценка векторов](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/store/sqlite/vectors.ts) | +| Модели, объединение и свежесть | [Сервис векторов](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/embeddings.ts), [каталог моделей](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/embeddings/models.ts) и [пул процессов](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/embeddings/pool.ts) | +| Контракт пакета доказательств | [Сервис доказательств](https://github.com/leemour/cli-messaging/blob/v0.148.2/src/services/evidence.ts) | + +Точная грамматика, поля, preset и пределы — в [справке по языку запросов](https://github.com/leemour/cli-messaging/blob/v0.148.2/docs/search/query-language.md). Подготовка и опции — на страницах архива [Telegram](./tg/archive.md) и [MAX](./max/archive.md). diff --git a/content/docs/search-playground.es.mdx b/content/docs/search-playground.es.mdx index d7b39ed..c81b7bb 100644 --- a/content/docs/search-playground.es.mdx +++ b/content/docs/search-playground.es.mdx @@ -11,4 +11,4 @@ Usa `has:` para buscar archivos, enlaces, fotos, vídeos, audio, mensajes de voz Consulta las guías de búsqueda de [Telegram](./tg/search.md) y [MAX](./max/search.md) para preparar el archivo y conocer los campos y las opciones. -Consulta cómo se conectan la búsqueda estricta, los grafos y los embeddings en la [guía de arquitectura de búsqueda](./search-architecture.mdx). +Consulta cómo se conectan la búsqueda estricta, los grafos y los embeddings en [cómo funciona la búsqueda](./search-architecture.mdx). diff --git a/content/docs/search-playground.mdx b/content/docs/search-playground.mdx index 62411a3..074a03f 100644 --- a/content/docs/search-playground.mdx +++ b/content/docs/search-playground.mdx @@ -11,4 +11,4 @@ Choose `has:` to try files, links, photos, video, audio, voice messages, sticker See the full [Telegram search guide](./tg/search.md) or [MAX search guide](./max/search.md) for archive setup, supported fields and command options. -Learn how strict search, conversation graphs and embeddings fit together in the [search architecture guide](./search-architecture.mdx). +Learn how strict search, conversation graphs and embeddings fit together on [how search works](./search-architecture.mdx). diff --git a/content/docs/search-playground.ru.mdx b/content/docs/search-playground.ru.mdx index f7742b1..46ed4d6 100644 --- a/content/docs/search-playground.ru.mdx +++ b/content/docs/search-playground.ru.mdx @@ -11,4 +11,4 @@ description: Попробуйте поиск сообщений Настройка архива, поля и параметры команд описаны в [руководстве по поиску Telegram](./tg/search.md) и [руководстве по поиску MAX](./max/search.md). -Как связаны строгий поиск, граф разговоров и embeddings — в [руководстве по архитектуре поиска](./search-architecture.mdx). +Как связаны строгий поиск, граф разговоров и embeddings — на странице [как устроен поиск](./search-architecture.mdx). diff --git a/docs/STRUCTURE.md b/docs/STRUCTURE.md index 92a6528..d21d812 100644 --- a/docs/STRUCTURE.md +++ b/docs/STRUCTURE.md @@ -26,6 +26,12 @@ not read in order. export, how long it lives. - `mcp.md` — the MCP server for clients without a terminal: connecting, what it can do, what is off until turned on. +- `search.md` — everyday message search: words, people, chats, dates, files, links, tags, saved + searches and counts. No internals. +- `topic-search.md` — conversations for a non-technical reader: building, embedding, searching by + meaning, freshness, and what a remote model sends. +- `query-language.md` — the search reference: fields, operators, presets, limits, the JSON answer. + The technical page on how search works is shared: `content/docs/search-architecture.mdx`. - `recipes.md` — whole tasks for an agent, each one copyable. - Pages only one tool has, named after what they cover — today `bot.md`, `groups.md`, `remote.md` (max).