Skip to content
Start free

Keep a Knowledge Base Current

Re-read a defined set of public pages with /extract/markdown, detect what actually changed with a content hash, and reprocess only those pages in your own pipeline.

A retrieval system is only as current as its last ingest. The usual failure is not that re-reading is hard, it is that re-embedding everything on every run is expensive, so it happens rarely, so the index goes stale.

This example re-reads a fixed source list, compares each page against the copy you stored last time, and hands back only the pages that changed. Embedding, chunking, and scheduling stay in your pipeline, where they belong.

import { createHash } from "node:crypto";
import { readFileSync, writeFileSync, existsSync } from "node:fs";
import Tabstack from "@tabstack/sdk";
const client = new Tabstack();
const STATE = "kb-state.json";
const SOURCES = [
"https://docs.tabstack.ai/guides/research",
"https://docs.tabstack.ai/guides/how-to-extract-json",
"https://docs.tabstack.ai/pricing",
];
function fingerprint(content: string): string {
return createHash("sha256").update(content).digest("hex");
}
function loadState(): Record<string, string> {
return existsSync(STATE) ? JSON.parse(readFileSync(STATE, "utf8")) : {};
}
/** Re-read each source and return only the ones whose content changed. */
async function refresh(sources: string[]) {
const state = loadState();
const changed: { url: string; content: string }[] = [];
for (const url of sources) {
// nocache matters here: the shared content cache is keyed by URL,
// effort and region, and a cached copy defeats the whole job.
const result = await client.extract.markdown({ url, nocache: true });
const digest = fingerprint(result.content);
if (state[url] === digest) {
console.log(`unchanged ${url}`);
continue;
}
console.log(`changed ${url}`);
changed.push({ url, content: result.content });
state[url] = digest;
}
writeFileSync(STATE, JSON.stringify(state, null, 2));
return changed;
}
for (const page of await refresh(SOURCES)) {
// Your pipeline owns what happens next: chunk, embed, upsert, delete.
console.log(`-> reprocess ${page.url} (${page.content.length} chars)`);
}

A run prints one line per source, then the pages your pipeline needs to reprocess:

unchanged https://docs.tabstack.ai/guides/research
changed https://docs.tabstack.ai/guides/how-to-extract-json
unchanged https://docs.tabstack.ai/pricing
-> reprocess https://docs.tabstack.ai/guides/how-to-extract-json (18422 chars)
  • One call per source, and nothing else. /extract/markdown is 10 credits, deterministic, and returns prose rather than markup. Re-reading a 40-page source set is a predictable cost you can budget.
  • nocache: true is required, not optional. Page content is cached by URL, effort, and region rather than by account, so without it a refresh can hand you the copy you already have. See Data Handling.
  • Hash the extracted markdown, not the HTML. A page whose nav, ads, or build hash changed will differ byte-for-byte in HTML while its content is identical. Extraction strips that, so the hash only moves when the prose moves.
  • Store the digest, not the document. The state file holds one hash per URL, so change detection costs nothing to keep around even for a large source set.

There is no scheduler in Tabstack. Use whatever already runs in your stack:

Terminal window
# Every night at 03:00
0 3 * * * cd /srv/kb && python refresh.py >> refresh.log 2>&1

A GitHub Actions cron, a Cloud Scheduler job, or a queue worker all work the same way. The only thing Tabstack provides is the current content of the pages you name.

  • Need typed fields rather than prose? Swap /extract/markdown for /extract/json with a schema, and hash the serialized object instead. Useful when you are tracking specific values rather than indexing text.
  • Source list that changes? Keep it in the same state file, or generate it from your own database. Tabstack does not crawl, so discovering new URLs is your step.
  • Pages that need a browser? Pass effort: 'max' for JavaScript-heavy sources. See Effort levels.
  • Want to know what changed, not just that it changed? Keep the previous markdown alongside the hash and diff the two. Prose diffs are readable, which is part of why markdown is the right storage format here.
Terminal window
npm install @tabstack/sdk

Set your API key before running:

Terminal window
export TABSTACK_API_KEY=your_api_key