diff --git a/CHANGELOG.md b/CHANGELOG.md
index a42f949..08dcf44 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -1,5 +1,23 @@
# Changelog / 更新记录
+## V0.1.1 — Unreleased
+
+English:
+
+* Added selectable AI Chat and API extraction modes.
+* Added local prompt generation for manual ChatGPT, Claude, Gemini, Qwen, and other AI chat workflows.
+* Added manual JSON import with malformed output handling, unknown-field checks, and the existing deterministic field validation.
+* Added manual extraction provenance without storing the raw pasted chat response.
+* Kept the existing BYOK API flow and provider region controls unchanged.
+
+中文:
+
+* 增加可选择的 AI Chat 与 API 两种 Extraction 模式。
+* 增加适用于 ChatGPT、Claude、Gemini、Qwen 和其他 AI Chat 的本地 Prompt 生成。
+* 增加手工 JSON 导入,并继续使用现有 Malformed Output、Unknown Fields 和字段 Validation。
+* 增加 Manual Extraction Provenance,原始粘贴 Chat Response 不写入持久化项目数据。
+* 保持现有 BYOK API 流程和 Provider Region 控制不变。
+
## V0.1.0 — 2026-09-19
English:
diff --git a/README.md b/README.md
index 81393e5..26b67c7 100644
--- a/README.md
+++ b/README.md
@@ -4,7 +4,7 @@ Local first AI assisted document extraction, validation and human review.
[简体中文](README.zh-CN.md)
-## Try V0.1.0 and send feedback
+## Try FlowExtract and send feedback
Live app: https://flowextract.edwardxie421.workers.dev
@@ -12,12 +12,21 @@ Testing FlowExtract with a real document? Please report bugs and real-world test
Do not include API keys, credentials, full provider responses, or sensitive document content in public issues. If a reproduction file is useful, use a public, synthetic, or sanitized example.
-FlowExtract is an open source browser application for turning PDFs and document images into structured, reviewable data. The first release is deliberately small: upload a document, define a schema, extract with your own AI provider key, validate results locally, correct questionable fields, and export JSON, CSV, or XLSX.
+FlowExtract is an open source browser application for turning PDFs and document images into structured, reviewable data. The current V0.1.x scope stays deliberately small: upload a document, define a schema, extract through an AI chat or your own AI provider key, validate results locally, correct questionable fields, and export JSON, CSV, or XLSX.
## V0.1 workflow
`Document -> Extract -> Validate -> Review -> Export`
+### Extraction modes
+
+FlowExtract supports two extraction paths that converge on the same Validation, Review, Final Value, Export, and local eval flow:
+
+* **AI Chat**: generate a document-specific prompt locally, copy it into ChatGPT, Claude, Gemini, Qwen, or another AI chat, then paste the JSON response back into FlowExtract. No API key is required. FlowExtract does not automate or read the user's AI chat session.
+* **API**: keep the existing automated BYOK flow using OpenAI, Anthropic, Gemini, or Qwen. Provider usage may consume API quota or incur provider charges.
+
+Manual AI Chat imports preserve the AI prediction and record the selected chat service as provenance. The raw pasted chat response remains page-local while editing and is not added to the project record.
+
The review step is the product focus. AI makes the first prediction. Deterministic rules identify missing, malformed, or out of range values. The user sees the source alongside the predicted and final values and corrects only what needs attention.
## V0.1 scope
@@ -39,7 +48,7 @@ Exports include JSON, CSV, and XLSX. Project state is stored in IndexedDB. Porta
FlowExtract has no application backend in V0.1.
-Documents are parsed in the browser. OCR runs locally with Tesseract.js. AI extraction sends the parsed document text directly from the browser to the provider selected by the user. FlowExtract does not proxy those requests through a FlowExtract server.
+Documents are parsed in the browser. OCR runs locally with Tesseract.js. In AI Chat mode, FlowExtract generates the prompt locally and the user decides when and where to paste it. In API mode, parsed document text is sent directly from the browser to the provider selected by the user. FlowExtract does not proxy either path through a FlowExtract server.
API keys are held only in React memory for the current page session. They are not written to IndexedDB, project backups, source code, or logs. Reloading the page clears the key.
diff --git a/README.zh-CN.md b/README.zh-CN.md
index d8a0a5b..c04982b 100644
--- a/README.zh-CN.md
+++ b/README.zh-CN.md
@@ -4,7 +4,7 @@ Local First 的 AI 辅助文档数据提取、验证与人工复核工具。
[English](README.md)
-## 试用 V0.1.0 与反馈
+## 试用 FlowExtract 与反馈
在线版本:https://flowextract.edwardxie421.workers.dev
@@ -12,12 +12,21 @@ Local First 的 AI 辅助文档数据提取、验证与人工复核工具。
请勿在公开 Issue 中提交 API Key、凭据、完整 Provider Response 或敏感文档正文。如需提供复现文件,请使用公开、虚构或已脱敏的示例。
-FlowExtract 是一个开源浏览器应用,用于把 PDF 和文档图片转换成可以人工复核的结构化数据。V0.1 刻意控制范围:上传文档,定义 Schema,使用自己的 AI Provider API Key 完成初始提取,在本地执行确定性验证,只修正可疑字段,然后导出 JSON、CSV 或 XLSX。
+FlowExtract 是一个开源浏览器应用,用于把 PDF 和文档图片转换成可以人工复核的结构化数据。当前 V0.1.x 继续控制范围:上传文档,定义 Schema,通过 AI Chat 或自己的 AI Provider API Key 完成初始提取,在本地执行确定性验证,只修正可疑字段,然后导出 JSON、CSV 或 XLSX。
## V0.1 工作流
`Document -> Extract -> Validate -> Review -> Export`
+### Extraction 模式
+
+FlowExtract 提供两种提取方式,两条路径最终都进入同一套 Validation、Review、Final Value、Export 和本地 Evals:
+
+* **AI Chat**:在本地生成包含当前文档和 Schema 的 Prompt,由用户自行复制到 ChatGPT、Claude、Gemini、Qwen 或其他 AI Chat,再把 JSON 结果粘贴回 FlowExtract。无需 API Key。FlowExtract 不会自动操作或读取用户的 AI Chat 会话。
+* **API**:保留现有自动化 BYOK 流程,支持 OpenAI、Anthropic、Gemini、Qwen。Provider 可能消耗 API 配额或产生费用。
+
+Manual AI Chat 导入后仍会保留 AI Prediction,并把用户选择的 Chat Service 记录为 Provenance。用户粘贴的原始 Chat Response 只在当前页面编辑状态中存在,不写入 Project Record。
+
Review 是产品重点。AI 先给出预测,确定性 Validation Rules 检查缺失值、格式错误和范围异常,用户同时查看来源文本、AI prediction 和 final value,只处理需要人工介入的字段。
## 当前 V0.1 范围
@@ -39,7 +48,7 @@ BYOK Provider 已建立 OpenAI、Anthropic、Gemini、Qwen(阿里云百炼 / M
V0.1 没有 FlowExtract 应用后端。
-文档在浏览器本地解析。OCR 使用 Tesseract.js 在本地执行。需要 AI 提取时,浏览器直接把解析后的文档文本发送给用户选择的 AI Provider,不经过 FlowExtract 自己的服务器。
+文档在浏览器本地解析。OCR 使用 Tesseract.js 在本地执行。AI Chat 模式在本地生成 Prompt,由用户自行决定何时、向哪个 AI Chat 粘贴。API 模式由浏览器把解析后的文档文本直接发送给用户选择的 AI Provider。两种方式都不经过 FlowExtract 自己的服务器。
API Key 只保存在当前页面的 React 内存状态中,不写入 IndexedDB、项目备份、源代码或日志。刷新页面后 Key 会消失。
diff --git a/SECURITY.md b/SECURITY.md
index 7114f93..d7f39f7 100644
--- a/SECURITY.md
+++ b/SECURITY.md
@@ -18,6 +18,12 @@ Production code should avoid logging document text, provider responses that may
生产代码应避免记录文档文本、可能包含文档信息的 Provider 响应以及认证 Header。
+## Manual AI Chat trust boundary / Manual AI Chat 信任边界
+
+AI Chat mode generates the extraction prompt entirely in the browser. FlowExtract does not open authenticated sessions, read cookies, automate third-party chat interfaces, or submit the prompt on the user's behalf. The user decides whether to copy the prompt and paste it into an external AI service. The generated prompt contains parsed document content, so the selected AI service's privacy and data-retention policies apply after the user pastes it there. The raw pasted chat response remains transient UI state; the project persists only parsed extraction data, validation, corrections, and manual service provenance.
+
+AI Chat 模式完全在浏览器中生成 Extraction Prompt。FlowExtract 不读取 Cookie、不接管登录 Session、不自动操作第三方聊天页面,也不代替用户发送 Prompt。用户自行决定是否复制 Prompt 并粘贴到外部 AI 服务。Prompt 包含解析后的文档内容,因此用户粘贴后应遵循对应 AI 服务的隐私和数据保留政策。原始粘贴 Chat Response 只作为临时 UI 状态存在,项目只持久化解析后的 Extraction 数据、Validation、Correction 和 Manual Service Provenance。
+
## Browser BYOK threat model / 浏览器 BYOK 威胁模型
V0.1 accepts a provider key at runtime because the project has no FlowExtract backend. The key is never committed or persisted, but any secret present in a browser page is accessible to that page runtime. OpenAI and Google recommend server-side handling for long-lived production keys. Use a dedicated low-limit key for V0.1 testing and revoke it if exposure is suspected.
diff --git a/docs/manual-ai-chat-mode.md b/docs/manual-ai-chat-mode.md
new file mode 100644
index 0000000..20cd68e
--- /dev/null
+++ b/docs/manual-ai-chat-mode.md
@@ -0,0 +1,51 @@
+# Manual AI Chat Extraction Mode
+
+Status: V0.1.1 release candidate.
+
+## Goal
+
+Lower the first-use barrier for people who already have access to an AI chat product and do not want to create or fund an API key.
+
+The core FlowExtract pipeline remains:
+
+`Document -> Extract -> Validate -> Review -> Export`
+
+## Modes
+
+### AI Chat
+
+1. Parse the document locally.
+2. Build a prompt locally from the parsed document text and current schema.
+3. Let the user copy the prompt.
+4. Let the user open ChatGPT, Claude, Gemini, Qwen, or another AI chat.
+5. Let the user paste the AI response back into FlowExtract.
+6. Parse and validate the response locally.
+7. Continue through the existing Review, correction, export, persistence, and eval flow.
+
+FlowExtract does not sign in to, automate, scrape, or read any AI chat account.
+
+### API
+
+The existing BYOK provider flow remains unchanged. Parsed document text is sent directly from the browser to the explicitly selected provider and region.
+
+## Data and provenance
+
+API extraction records `extractionMode: "api"` plus provider, model, optional region, timestamp, predictions, validation, and corrections.
+
+Manual extraction records `extractionMode: "manual"` plus the selected chat service, timestamp, predictions, validation, and corrections.
+
+Legacy V0.1.0 extraction records without `extractionMode` are treated as API records for UI restoration.
+
+The raw pasted AI chat response is transient UI state and is not written to the project record, IndexedDB, or project backup.
+
+## Validation
+
+Manual imports reuse the existing deterministic validation engine. They keep the same behavior for Required, Type, Regex, Minimum, Maximum, strict YYYY-MM-DD dates, Unknown Fields, and Malformed AI Output.
+
+Malformed pasted output creates an extraction result with a global validation issue so the failure is visible in Review.
+
+## Security boundary
+
+The generated prompt contains parsed document text. Copying it does not transmit data. Once the user pastes it into an external AI service, that service's privacy, retention, billing, and account policies apply.
+
+No browser session reuse, cookie access, hidden API, web scraping, or cross-service automation is part of this mode.
diff --git a/package-lock.json b/package-lock.json
index 83386db..20ae41a 100644
--- a/package-lock.json
+++ b/package-lock.json
@@ -1,12 +1,12 @@
{
"name": "flowextract",
- "version": "0.1.0",
+ "version": "0.1.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "flowextract",
- "version": "0.1.0",
+ "version": "0.1.1",
"dependencies": {
"exceljs": "^4.4.0",
"pdfjs-dist": "^5.4.149",
diff --git a/package.json b/package.json
index b69281c..aadbf51 100644
--- a/package.json
+++ b/package.json
@@ -1,6 +1,6 @@
{
"name": "flowextract",
- "version": "0.1.0",
+ "version": "0.1.1",
"private": true,
"type": "module",
"description": "Local first AI assisted document extraction, validation and human review.",
diff --git a/src/App.test.tsx b/src/App.test.tsx
index 8e06d4b..58b9c4b 100644
--- a/src/App.test.tsx
+++ b/src/App.test.tsx
@@ -1,4 +1,4 @@
-import { act, cleanup, render, screen, waitFor } from '@testing-library/react';
+import { act, cleanup, fireEvent, render, screen, waitFor } from '@testing-library/react';
import { afterEach, describe, expect, it, vi } from 'vitest';
import App from './App';
import * as projectStore from './features/persistence/projectStore';
@@ -22,12 +22,58 @@ describe('FlowExtract workspace', () => {
expect(screen.getByText(/2\. Schema/i)).toBeInTheDocument();
expect(screen.getByText(/3\. AI Extraction/i)).toBeInTheDocument();
expect(screen.getByText(/4\. Review/i)).toBeInTheDocument();
- expect(screen.getByText('v0.1.0')).toBeInTheDocument();
- expect(screen.getByRole('option', { name: /Qwen \(Alibaba Cloud\) — Verified in Beijing/i })).toBeInTheDocument();
+ expect(screen.getByText('v0.1.1')).toBeInTheDocument();
+ expect(screen.getByRole('button', { name: /AI Chat/i })).toHaveAttribute('aria-pressed', 'true');
+ expect(screen.getByRole('option', { name: 'ChatGPT' })).toBeInTheDocument();
+ fireEvent.click(screen.getByRole('button', { name: /^API/i }));
+ expect(screen.getByRole('option', { name: /Qwen \(Alibaba Cloud\) - Verified in Beijing/i })).toBeInTheDocument();
+ expect(screen.getByPlaceholderText('Stored in memory only')).toHaveValue('');
expect(screen.getByRole('link', { name: 'Feedback' })).toHaveAttribute('href', 'https://github.com/edwardsage419/FlowExtract/issues/new/choose');
});
+
+ it('imports pasted AI chat JSON into the existing validation and review flow', async () => {
+ const source: ProjectRecord = {
+ id: 'manual-source',
+ name: 'Manual chat invoice',
+ updatedAt: '2026-09-19T00:00:00.000Z',
+ document: {
+ id: 'doc-manual',
+ name: 'invoice.pdf',
+ mimeType: 'application/pdf',
+ size: 100,
+ createdAt: '2026-09-19T00:00:00.000Z',
+ text: 'Invoice INV-42 total 1333.80',
+ pages: ['Invoice INV-42 total 1333.80'],
+ sourceKind: 'pdf',
+ ocrUsed: false,
+ },
+ schema: {
+ id: 'schema-manual',
+ name: 'Invoice',
+ updatedAt: '2026-09-19T00:00:00.000Z',
+ fields: [{ id: 'amount-field', name: 'Amount', key: 'amount', type: 'number', required: true, description: '', rules: {} }],
+ },
+ };
+ vi.spyOn(projectStore, 'listProjects').mockResolvedValue([source]);
+ vi.spyOn(projectStore, 'saveProject').mockResolvedValue();
+ render();
+
+ await waitFor(() => expect(screen.getByRole('textbox', { name: /Project/i })).toHaveValue('Manual chat invoice'));
+ expect(screen.getByRole('button', { name: /AI Chat/i })).toHaveAttribute('aria-pressed', 'true');
+
+ const fence = String.fromCharCode(96).repeat(3);
+ fireEvent.change(screen.getByLabelText('AI chat response'), {
+ target: { value: fence + 'json\n{"amount":1333.8}\n' + fence },
+ });
+ fireEvent.click(screen.getByRole('button', { name: 'Import & validate' }));
+
+ await waitFor(() => expect(screen.getByDisplayValue('1333.8')).toBeInTheDocument());
+ expect(screen.getByText('1333.8')).toBeInTheDocument();
+ expect(screen.getByText(/0 issues/i)).toBeInTheDocument();
+ });
+
it('does not autosave a blank project before startup hydration completes', async () => {
vi.useFakeTimers();
let resolveProjects!: (projects: ProjectRecord[]) => void;
@@ -95,6 +141,7 @@ describe('FlowExtract workspace', () => {
expect(screen.getByDisplayValue('1333.81')).toBeInTheDocument();
expect(screen.getByText('1333.8')).toBeInTheDocument();
expect(screen.getAllByText(/corrected/i).length).toBeGreaterThan(0);
+ expect(screen.getByRole('button', { name: /^API/i })).toHaveAttribute('aria-pressed', 'true');
expect(screen.getByPlaceholderText('Stored in memory only')).toHaveValue('');
});
});
diff --git a/src/App.tsx b/src/App.tsx
index 34c0422..00d34fd 100644
--- a/src/App.tsx
+++ b/src/App.tsx
@@ -7,6 +7,8 @@ import { parseDocument } from './features/documents/parseDocument';
import { computeMetrics } from './features/evals/metrics';
import { toCsv, toJson, toXlsx } from './features/export/exporters';
import { runExtraction } from './features/extraction/extract';
+import { buildManualExtractionPrompt, importManualExtraction } from './features/extraction/manual';
+import type { ExtractionMode, ManualAIService } from './features/extraction/types';
import { exportProjectBackup, importProjectBackup } from './features/persistence/backup';
import { listProjects, loadProject, saveProject } from './features/persistence/projectStore';
import type { ProjectRecord } from './features/project/types';
@@ -39,6 +41,9 @@ function downloadText(filename: string, text: string, mime: string) { downloadBl
export default function App() {
const [project, setProject] = useState(() => newProject());
+ const [extractionMode, setExtractionMode] = useState('manual');
+ const [manualService, setManualService] = useState('chatgpt');
+ const [manualResponse, setManualResponse] = useState('');
const [provider, setProvider] = useState('openai');
const [model, setModel] = useState(DEFAULT_MODELS.openai);
const [apiKey, setApiKey] = useState('');
@@ -53,6 +58,36 @@ export default function App() {
const restoreRef = useRef(null);
const metrics = useMemo(() => project.extraction ? computeMetrics(project.extraction.fields) : null, [project.extraction]);
+ const manualPrompt = useMemo(() => {
+ if (!project.document) return '';
+ try {
+ return buildManualExtractionPrompt(project.document, project.schema);
+ } catch {
+ return '';
+ }
+ }, [project.document, project.schema]);
+
+ function restoreExtractionUi(source?: ProjectRecord) {
+ const extraction = source?.extraction;
+ setApiKey('');
+ setManualResponse('');
+ if (!extraction) {
+ setExtractionMode('manual');
+ setManualService('chatgpt');
+ return;
+ }
+ if (extraction.extractionMode === 'manual') {
+ setExtractionMode('manual');
+ setManualService(extraction.manualService ?? 'other');
+ return;
+ }
+ setExtractionMode('api');
+ if (extraction.provider) {
+ setProvider(extraction.provider);
+ setModel(extraction.model || DEFAULT_MODELS[extraction.provider]);
+ if (extraction.provider === 'qwen' && extraction.providerRegion) setQwenRegion(extraction.providerRegion);
+ }
+ }
useEffect(() => {
let cancelled = false;
@@ -64,7 +99,7 @@ export default function App() {
if (latest) {
setProject(latest);
setPreviewUrl(null);
- setApiKey('');
+ restoreExtractionUi(latest);
}
setHydrated(true);
})
@@ -84,6 +119,7 @@ export default function App() {
const touch = (next: ProjectRecord): ProjectRecord => ({ ...next, updatedAt: now() });
const updateSchema = (updater: (schema: ProjectRecord['schema']) => ProjectRecord['schema']) => {
+ setManualResponse('');
setProject((current) => touch({ ...current, schema: { ...updater(current.schema), updatedAt: now() }, extraction: undefined }));
};
@@ -94,6 +130,7 @@ export default function App() {
setPreviewUrl(URL.createObjectURL(file));
const documentRecord = await parseDocument(file, { ocrLanguage, onProgress: setProgress });
setProject((current) => touch({ ...current, document: documentRecord, extraction: undefined }));
+ setManualResponse('');
setProgress(documentRecord.ocrUsed ? 'Local OCR complete.' : 'PDF text extraction complete.');
} catch (cause) {
setError(cause instanceof Error ? cause.message : 'Unable to read document.'); setProgress('');
@@ -110,6 +147,22 @@ export default function App() {
finally { setBusy(null); }
}
+ function handleManualImport() {
+ if (!project.document) return;
+ setError('');
+ try {
+ const extraction = importManualExtraction({
+ document: project.document,
+ schema: project.schema,
+ service: manualService,
+ rawResponse: manualResponse,
+ });
+ setProject((current) => touch({ ...current, extraction }));
+ } catch (cause) {
+ setError(cause instanceof Error ? cause.message : 'Unable to import AI chat response.');
+ }
+ }
+
function correct(key: string, value: string) {
if (!project.extraction) return;
const fields = applyCorrection(project.schema, project.extraction.fields, key, value);
@@ -130,21 +183,33 @@ export default function App() {
function exportBackup() { downloadText(`${project.name.replace(/[^a-z0-9]+/gi, '-').toLowerCase() || 'flowextract'}-backup.json`, exportProjectBackup(project), 'application/json'); }
async function restoreBackup(file: File) {
- try { const restored = importProjectBackup(await file.text()); setProject(restored); setPreviewUrl(null); setApiKey(''); setError(''); }
- catch (cause) { setError(cause instanceof Error ? cause.message : 'Unable to restore backup.'); }
+ try {
+ const restored = importProjectBackup(await file.text());
+ setProject(restored);
+ setPreviewUrl(null);
+ restoreExtractionUi(restored);
+ setError('');
+ } catch (cause) { setError(cause instanceof Error ? cause.message : 'Unable to restore backup.'); }
+ }
+ async function openRecent(id: string) {
+ const stored = await loadProject(id);
+ if (stored) {
+ setProject(stored);
+ setPreviewUrl(null);
+ restoreExtractionUi(stored);
+ }
}
- async function openRecent(id: string) { const stored = await loadProject(id); if (stored) { setProject(stored); setPreviewUrl(null); setApiKey(''); } }
return (
-
FX
FlowExtract
v0.1.0
Local first AI assisted document extraction, validation and human review.
+
FX
FlowExtract
v0.1.1
Local first AI assisted document extraction, validation and human review.