@lacspace/sourcewatch
Re-verify that official source pages still back the facts you published: fetch a URL (HTML, plain text or PDF), extract its visible text, normalise it (NFC, Devanagari digits to ASCII, zero-width and soft-hyphen removal, unified spaces and dashes) and check expected strings, whole numbers ("1145" never matches inside "11450") or RegExps. Pure-JS PDF text extraction (object streams, FlateDecode/LZW/ASCII85/ASCIIHex with PNG predictors, ToUnicode CMaps, Differences encodings, TJ kerning, page-tree order). Detects placeholder pages ("this is test", default server pages, coming soon, suspended accounts), classifies failures (timeout, tls with chain/expired/self_signed/hostname kind, dns, network, http_4xx, http_5xx, not_found_text, placeholder_page, too_large, unparseable), streams bodies with a byte cap, retries transient errors, hashes the normalised text for change detection, and runs batches with a concurrency limit and one-request-per-host politeness. check(), checkAll(), summarize(), pdfText(), extractText(), normalise(), isPlaceholder(). Injectable fetch, zero dependencies, isomorphic (Node 18+, browsers, edge).
npm i @lacspace/sourcewatchUsage
import { check, checkAll, summarize } from "@lacspace/sourcewatch";
const r = await check({ id: "water-helpline", url: "https://nwc.gov.np", expect: "1145" });
// { ok: true, status: 200, kind: "html", found: true, matched: ["1145"], missing: [],
// snippet: "…प्रेष विज्ञप्ति 1145 मा सम्पर्कका लागि अनुरोध…", contentHash: "590bd7c96affa346", ms: 1222, … }
const results = await checkAll([
{ id: "short-codes", url: "https://nta.gov.np/uploads/contents/National%20Numbering%20Allocation%20Plan.pdf", expect: ["100", "101", "102"] },
{ id: "eoc", url: "http://neoc.gov.np", expect: "1149" }, // → placeholder_page ("this is test")
{ id: "results", url: "https://neb.gov.np", expect: /results?/i, prevHash: lastRun.results },
], { concurrency: 4 });
summarize(results); // { total: 3, ok: 2, failed: 1, changed: 0, byError: { placeholder_page: 1 } }Exports 16
DEFAULT_USER_AGENTcheckcheckAllclassifyErrordecodeEntitiesextractTextfnv1a64htmlToTextisPdfisPlaceholdermatchExpectnormalisepdfTextplaceholderReasonsummarizetlsKindOfKeywords
More in Automation Kit
Drive every describe()-capable @lacspace package through one interface — merge descriptors into a single command catalogue, render it token-light for an LLM prompt, validate an AI-written plan against the command schemas (unknown commands, required inputs, types/enums, forward references, duplicate ids), and execute it step by step with $steps.<id>.output references, parallel groups, per-step timeouts, a whole-plan budget, optional steps, when-conditions and a dry run. The AI decides; the packages do the work. Zero deps.
@lacspace/extractiveTurn a full article into a compact brief for an AI writer — extractive TextRank summary, key-facts (numbers, money, percentages, dates, named entities), headline candidates and key phrases/hashtags, for English and Nepali (sentence splitting on the danda '।', decimal-safe). Deterministic, no LLM; reuses @lacspace/keyphrase and @lacspace/factcheck-lite. Exposes a describe() command schema so an AI 'conductor' can drive it.
@lacspace/hookwriterDeterministic platform copy from facts — hooks, titles, captions, CTAs and descriptions in many styles (question, number-led, what-it-means, contrast, breaking, how-to, quote…), English and Nepali, trimmed to each platform's character limit (YouTube/IG/TikTok/FB/X/Threads/LinkedIn/Telegram). Fills ONLY the facts you pass (never invents claims) and drops sensational phrasing; AI optional, only to fill slots. Exposes a describe() command schema for an AI 'conductor'. Zero dependencies.
@lacspace/postbanditA tiny Thompson-sampling multi-armed bandit to learn the best option per platform — posting time, format, thumbnail or title variant — from engagement rewards. Beta posteriors per arm, deterministic when seeded, JSON-persistable; plugs into postplan. Zero deps.
@lacspace/commentguardComment moderation for English, romanized Nepali and Devanagari — scores spam, abuse, hate, doxxing and link-spam and returns an allow / review / flag / hide action, with PII detection (Nepali phone, email, URLs/shorteners) and FAQ auto-reply suggestions. Deterministic rules + small, extendable lexicons; a borderline flag marks gray-zone comments for an optional AI second opinion. Exposes a describe() command schema. Zero dependencies.
@lacspace/gov-noticesParse Nepali government and university notice boards into clean, dated, tagged items for exam and results feeds. Site adapters for PSC (Lok Sewa), NEB, SEE (OCE Sanothimi), TSC, MEC, CTEVT, Nepal Engineering Council and DoTM, plus a generic notice-list finder for any other board. Each notice has a title, Devanagari-aware language flag, AD + Bikram Sambat dates (parses २०८२/०६/१८, 2082-06-18, २०८३ असोज १८, Asoj 18 2082, Sep 15 2026), absolute URL, typed PDF/image/doc attachments and auto tags (result, exam, schedule, admit-card, vacancy, syllabus). parseNotices(), fetchNotices() with conditional GET (ETag / Last-Modified), dedupe(), newSince() for 'results out' alerts, parseBsDate()/parseDate(). Polite: one request per call. Zero runtime dependencies, isomorphic, no AI, no keys.