Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
72 changes: 53 additions & 19 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# 🔥 Firecrawl CLI

Command-line interface for Firecrawl. Search, scrape, interact, crawl, map, search research papers and developer sources, and run agent jobs directly from your terminal.
Command-line interface for Firecrawl. Search, scrape, interact, crawl, map, search research papers, developer sources, and government sources, and run agent jobs directly from your terminal.

## Installation

Expand Down Expand Up @@ -299,7 +299,7 @@ firecrawl search "landscape photography" --sources images
# Multiple sources
firecrawl search "machine learning" --sources web,news,images

# Filter by category (research-affiliated websites, PDFs, developer index)
# Filter by category (research-affiliated websites, PDFs, developer index, gov index)
firecrawl search "transformer architecture" --categories research
firecrawl search "machine learning" --categories pdf,research

Expand All @@ -309,6 +309,9 @@ firecrawl search "machine learning" --categories pdf,research
# Developer search: public repositories, GitHub issues, merged PRs, READMEs, and docs
firecrawl search "axum middleware ordering" --categories developer

# Government search: US government sources (cannot be combined with other categories)
firecrawl search "California data breach notification statute" --categories gov

# Time-based search
firecrawl search "AI announcements" --tbs qdr:d # Past day
firecrawl search "tech news" --tbs qdr:w # Past week
Expand All @@ -327,23 +330,23 @@ firecrawl search "AI data tools"

#### Search Options

| Option | Description |
| ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--limit <n>` | Maximum results (default: 5, max: 100) |
| `--sources <sources>` | Comma-separated: `web`, `images`, `news` (default: web) |
| `--categories <categories>` | Comma-separated: `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer` |
| `--tbs <value>` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) |
| `--location <location>` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") |
| `--country <code>` | ISO country code (default: US) |
| `--timeout <ms>` | Timeout in milliseconds (default: 60000) |
| `--highlights` | Query-relevant highlights for web and news when available (default) |
| `--no-highlights` | Keep the original search snippets |
| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints |
| `--scrape` | Enable scraping of search results |
| `--scrape-formats <formats>` | Scrape formats when `--scrape` enabled (default: markdown) |
| `--only-main-content` | Include only main content when scraping (default: true) |
| `-o, --output <path>` | Save to file |
| `--json` | Output as compact JSON |
| Option | Description |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--limit <n>` | Maximum results (default: 5, max: 100) |
| `--sources <sources>` | Comma-separated: `web`, `images`, `news` (default: web) |
| `--categories <categories>` | Comma-separated: `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer`, `gov` (cannot be combined with other categories) |
| `--tbs <value>` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) |
| `--location <location>` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") |
| `--country <code>` | ISO country code (default: US) |
| `--timeout <ms>` | Timeout in milliseconds (default: 60000) |
| `--highlights` | Query-relevant highlights for web and news when available (default) |
| `--no-highlights` | Keep the original search snippets |
| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints |
| `--scrape` | Enable scraping of search results |
| `--scrape-formats <formats>` | Scrape formats when `--scrape` enabled (default: markdown) |
| `--only-main-content` | Include only main content when scraping (default: true) |
| `-o, --output <path>` | Save to file |
| `--json` | Output as compact JSON |

#### Examples

Expand Down Expand Up @@ -409,6 +412,37 @@ firecrawl developer "tokio select cancellation safety" --json -o results.json

---

### `gov` - Search the Firecrawl Government Index

Search the Government Index: primary law and regulatory material from US federal, state, and local government sources, including statutes, regulations, codes, court opinions, and other government publications.

For the request and response schema, see the [Government Index REST API](https://docs.firecrawl.dev/features/gov).

```bash
firecrawl gov "food labeling requirements for allergens"
```

#### Options

| Option | Description |
| --------------------- | ----------------------------------------- |
| `--limit <n>` | Number of results (default: 10, max: 100) |
| `-o, --output <path>` | Save to file |
| `--json` | Output the raw response as JSON |
| `--pretty` | Pretty print JSON output |

#### Examples

```bash
# Find state statutes on a topic
firecrawl gov "California data breach notification statute" --limit 10

# Keep the raw response
firecrawl gov "FDA food labeling regulations" --json -o results.json
```

---

### `research` - Search research papers

Search Firecrawl's research paper index: roughly 43M abstracts, around 90% biomedical (PubMed, bioRxiv, medRxiv) plus arXiv. Use this for biomedical, clinical, and scientific literature rather than scraping PubMed, bioRxiv, or Google Scholar by hand.
Expand Down
2 changes: 1 addition & 1 deletion skills/firecrawl-search/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ firecrawl search "your query" --sources news --tbs qdr:d -o .firecrawl/news.json

Use `firecrawl search --help` for search options, `firecrawl list --help` for contract browsing, and `firecrawl scrape --help` for execution options.

`--categories developer` searches an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md).
`--categories developer` searches an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. `--categories gov` searches US federal, state, and local government legal and regulatory sources and cannot be combined with other categories. `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md).

**Done when:** relevant results have been inspected, per-call errors and empty results have been checked, the request has been answered with source links, and feedback is sent within the time window unless opted out.

Expand Down
10 changes: 10 additions & 0 deletions src/__tests__/cli-argv.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,16 @@ describe('CLI argv parsing', () => {
expect(result.stderr).not.toContain('unknown command');
});

testWithBuiltCli('parses the gov command and shows its help', () => {
const result = spawnSync(process.execPath, [cliPath, 'gov', '--help'], {
cwd: process.cwd(),
encoding: 'utf8',
});

expect(result.status).toBe(0);
expect(result.stdout).toContain('Usage: firecrawl gov');
});

testWithBuiltCli(
'describes default search highlights and public developer coverage',
() => {
Expand Down
186 changes: 186 additions & 0 deletions src/__tests__/commands/gov.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,186 @@
/**
* Tests for gov command
*/

import { describe, it, expect, vi, beforeEach, afterEach } from 'vitest';
import { handleGovSearchCommand } from '../../commands/gov';
import { getClient, isKeylessMode } from '../../utils/client';
import { initializeConfig } from '../../utils/config';
import { writeOutput } from '../../utils/output';
import { setupTest, teardownTest } from '../utils/mock-client';

vi.mock('../../utils/output', () => ({ writeOutput: vi.fn() }));

vi.mock('../../utils/client', async () => {
const actual = await vi.importActual('../../utils/client');
return {
...actual,
getClient: vi.fn(),
isKeylessMode: vi.fn(() => false),
};
});

describe('handleGovSearchCommand', () => {
let mockHttpGet: ReturnType<typeof vi.fn>;

// Wrap a payload in the axios envelope returned by `client.http.get`.
const mockGovResponse = (web: any[]) => ({
data: { success: true, data: { web } },
});

const sampleResult = {
url: 'https://www.ecfr.gov/current/title-21/chapter-I/subchapter-B/part-101',
title: '21 CFR Part 101 -- Food Labeling',
description: 'Food labeling requirements for packaged foods.',
position: 1,
};

beforeEach(() => {
setupTest();
initializeConfig({
apiKey: 'test-api-key',
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
apiUrl: 'https://api.firecrawl.dev',
});

mockHttpGet = vi.fn();
vi.mocked(getClient).mockReturnValue({
http: { get: mockHttpGet },
} as any);
});

afterEach(() => {
teardownTest();
vi.clearAllMocks();
vi.unstubAllGlobals();
});

describe('API call generation', () => {
it.each([
Comment thread
capy-ai[bot] marked this conversation as resolved.
[{}, '/v2/search/gov?query=food+labeling&integration=cli'],
[{ k: 5 }, '/v2/search/gov?query=food+labeling&k=5&integration=cli'],
])('calls /v2/search/gov with %o', async (extra, expectedUrl) => {
mockHttpGet.mockResolvedValue(mockGovResponse([sampleResult]));

await handleGovSearchCommand({
query: 'food labeling',
...extra,
});

expect(mockHttpGet).toHaveBeenCalledTimes(1);
expect(mockHttpGet).toHaveBeenCalledWith(expectedUrl);
});
});

describe('output', () => {
it('renders numbered title, url, and description blocks', async () => {
mockHttpGet.mockResolvedValue(
mockGovResponse([
sampleResult,
{
url: 'https://www.ecfr.gov/current/title-21/part-102',
title: '21 CFR Part 102',
position: 2,
},
])
);

await handleGovSearchCommand({ query: 'food labeling' });

const [content] = vi.mocked(writeOutput).mock.calls[0];
expect(content).toBe(
[
'## 1. 21 CFR Part 101 -- Food Labeling',
sampleResult.url,
'Food labeling requirements for packaged foods.',
'',
'## 2. 21 CFR Part 102',
'https://www.ecfr.gov/current/title-21/part-102',
].join('\n')
);
});

it('prints a placeholder when the response has no data', async () => {
mockHttpGet.mockResolvedValue({ data: { success: true } });

await handleGovSearchCommand({ query: 'no hits' });

const [content] = vi.mocked(writeOutput).mock.calls[0];
expect(content).toBe('(no results)');
});

it('outputs the raw response as JSON with --json', async () => {
mockHttpGet.mockResolvedValue(mockGovResponse([sampleResult]));

await handleGovSearchCommand({
query: 'food labeling',
json: true,
});

const [content] = vi.mocked(writeOutput).mock.calls[0] as [string];
expect(JSON.parse(content)).toEqual({
success: true,
data: { web: [sampleResult] },
});
});
});

describe('keyless mode', () => {
it('calls the endpoint directly and renders the results', async () => {
vi.mocked(isKeylessMode).mockReturnValueOnce(true);
const fetchMock = vi.fn(
async (_url: string, _init?: RequestInit) =>
new Response(
JSON.stringify({ success: true, data: { web: [sampleResult] } }),
{ status: 200 }
)
);
vi.stubGlobal('fetch', fetchMock);

await handleGovSearchCommand({ query: 'food labeling' });

expect(mockHttpGet).not.toHaveBeenCalled();
expect(fetchMock).toHaveBeenCalledWith(
'https://api.firecrawl.dev/v2/search/gov?query=food+labeling&integration=cli',
expect.objectContaining({ method: 'GET' })
);
expect(vi.mocked(writeOutput).mock.calls[0][0]).toContain(
sampleResult.title
);
});
});

describe('error handling', () => {
it.each([
[
'the response reports a failure',
() =>
mockHttpGet.mockResolvedValue({
data: { success: false, error: 'Search failed' },
}),
'Search failed',
],
[
'the request fails',
() => mockHttpGet.mockRejectedValue(new Error('boom')),
'boom',
],
])('exits with code 1 when %s', async (_label, arrange, message) => {
arrange();
const exitSpy = vi
.spyOn(process, 'exit')
.mockImplementation((() => undefined) as any);
const errorSpy = vi
.spyOn(console, 'error')
.mockImplementation(() => undefined);

await handleGovSearchCommand({ query: 'test' });

expect(errorSpy).toHaveBeenCalledWith('Error:', message);
expect(exitSpy).toHaveBeenCalledWith(1);
expect(writeOutput).not.toHaveBeenCalled();

exitSpy.mockRestore();
errorSpy.mockRestore();
});
});
});
Loading
Loading