Line-aware OCR and data-loss prevention for scanned documents.
Watch a folder. Read every page. Quarantine sensitive documents. Keep the evidence.
DLPDuck watches a PDF drop folder, extracts native text and OCR page by page, applies position-aware DLP rules, and routes each document to an archive or a quarantine. Decisions go into a hash-chained audit trail; the web console gives operators one place to investigate hits, search text, retry failures, reassess documents under new rules, approve releases, and purge retained content.
| Built for | What DLPDuck does |
|---|---|
| Scanners and MFPs | Waits for stable PDF, TIFF, JPEG or PNG files and their companion metadata before claiming them |
| Mixed PDFs | Uses native text where it is trustworthy and OCR on sparse or image-bearing pages |
| Position-sensitive policy | Matches by page, line, document scope, and configurable line ranges |
| Sensitive findings | Stores masked values and keyed correlation digests, never the raw match |
| Operational review | Provides quarantine, a Needs attention queue, search, reassessment, and release workflows |
| Evidential history | Keeps append-only assessments and a verifiable, hash-chained audit log |
Note
DLPDuck is pre-1.0. Pin a version in production.
- How it works
- Install
- Docker
- Kubernetes
- Quick start
- Configuration
- Writing rules
- Storage, purge, and maintenance
- Audit trail
- Admin console
- CLI reference
- Writing plugins
- The watcher waits until a PDF (or a TIFF/JPEG/PNG scan, converted losslessly to PDF) stops changing, then moves it into staging.
- Each page uses native text when suitable and OCR when it is sparse; on a page with both, native text is kept and OCR adds only what is inside the images. A measurably blank page is not treated as a failed one.
- Rules inspect lines or the complete document and mask every recorded match.
- Enrichment plugins can add routing context before the final decision.
- Clean documents enter the archive; matches and incomplete extraction enter quarantine.
- DLPDuck writes the PDF, searchable content, assessment index, and audit event before notifying emit plugins.
A file that cannot be assessed never passes as clean. Findings contain masked text and a keyed correlation digest, never the raw matched value.
Requires Python 3.12+.
cd dlpduck # from a source checkout
uv sync # or: pip install -e .Two environment variables are required before anything runs:
export DLPDUCK_HMAC_KEY="$(openssl rand -hex 32)" # correlation digests
export DLPDUCK_SESSION_SECRET="$(openssl rand -hex 32)" # console cookie signingDLPDUCK_HMAC_KEY must be stable for the life of a deployment and must never be
written into config, the archive, or the audit trail.
export DLPDUCK_LOG_LEVEL=DEBUG # DEBUG | INFO (default) | WARNING | ERROR | CRITICALDEBUG is safe to turn on in production — it traces the pipeline (claim,
extract, scan, disposition, commit), rule match counts, and plugin runs, but
never a raw document value. DLPDUCK_LOG_LEVEL set to something unrecognised
refuses to start rather than silently falling back to a default.
Actual document content — full extracted text, and a rule's raw (unmasked) match — is behind a second, independent switch:
export DLPDUCK_TRACE_CONTENT_OUTPUT=trueThis does nothing unless DLPDUCK_LOG_LEVEL=DEBUG is also set — either one
alone is inert. Logs are typically the least access-controlled, longest
-retained thing in a deployment, which is exactly what masking.py and the
whole two-store split exist to keep raw document content out of; this flag
is for a deliberate, temporary debugging session, not a standing setting.
docker pull c0dewhacker/dlpduckSet DLPDUCK_RUN_BOTH=true to run the watcher and console together:
docker run -d --name dlpduck \
-e DLPDUCK_HMAC_KEY="$(openssl rand -hex 32)" \
-e DLPDUCK_SESSION_SECRET="$(openssl rand -hex 32)" \
-e DLPDUCK_RUN_BOTH=true \
-v /srv/dlpduck/config.yaml:/etc/dlpduck/config.yaml:ro \
-v /srv/dlpduck/drops:/srv/drops \
-v /srv/dlpduck/archive:/srv/archive \
-v /srv/dlpduck/quarantine:/srv/quarantine \
-v /srv/dlpduck/work:/srv/work \
-v /srv/dlpduck/audit:/srv/audit \
-p 8080:8080 \
c0dewhacker/dlpduck --config /etc/dlpduck/config.yamlSet audit.path: /srv/audit in the config for the separate audit mount. Without
it, audit data remains under the mounted work_dir/audit. The endpoints
/health/live and /health/ready provide container liveness and watcher
readiness checks. Omit DLPDUCK_RUN_BOTH and pass run or console run to run
one role. Container paths in the config must match the mounted paths, which must
be writable by the image's non-root user.
A Helm chart in charts/dlpduck runs the watcher and console
together with separate persistent claims for incoming files, archive,
quarantine, working data and the audit trail:
helm upgrade --install dlpduck charts/dlpduck \
--namespace dlpduck --create-namespace -f my-values.yamlSee the deployment guide for secrets, local or OIDC login, storage, ingress, TLS, upgrades and recovery.
mkdir -p /srv/dlpduck/{drops,archive,quarantine,work}
cat > config.yaml <<'YAML'
source:
name: mfp-3f
path: /srv/dlpduck/drops
metadata_format: none
destination:
archive: /srv/dlpduck/archive
quarantine: /srv/dlpduck/quarantine
work_dir: /srv/dlpduck/work
dlp:
rules:
- include: builtin:default.yaml
YAML
dlpduck validate-config --config config.yaml # compiles every regex, checks paths, warns
dlpduck run --config config.yaml # start the watcherIn another shell:
dlpduck console run --config config.yaml # http://127.0.0.1:8080Before dropping real documents, dry-run one:
dlpduck scan some.pdf --config config.yaml # prints lines + what would hitA minimal config is source, destination, and dlp.rules; everything else has
a default. The full shape:
Environment variables named DLPDUCK__SECTION__FIELD override YAML before
validation. For example, DLPDUCK__CONSOLE__BIND=0.0.0.0:8080 overrides
console.bind; lists and mappings accept JSON or inline YAML. The HMAC and
session values remain in DLPDUCK_HMAC_KEY and DLPDUCK_SESSION_SECRET, so
they never need to appear in the configuration file.
version: 2 # the config format; a mismatch is refused, not guessed at
umask: "0077" # owner-only for everything written; null to inherit
source:
name: mfp-3f # recorded on every job
path: /srv/dlpduck/drops
pdf_suffix: .pdf # case-insensitive
image_suffixes: [.tif, .tiff, .jpg, .jpeg, .png] # converted to PDF at claim
metadata_format: none # xml | json | text | none
metadata_suffix: .xml # companion file: scan.pdf + scan.xml
metadata_fields: [] # ALLOWLIST — see the note below
poll_seconds: 5.0
stability_polls: 2 # size must hold this many polls before claiming
metadata_grace_polls: 3 # extra polls to wait for a companion that's en route
limits:
max_bytes: 209715200 # 200 MB
max_pages: 500
rule_budget_ms: 2000 # per rule, per document
max_metadata_bytes: 1048576 # 1 MB — companion files are bounded too
max_image_pixels: 150000000 # per image frame, checked before decoding
extraction:
dpi: 150 # OCR raster resolution
native_min_chars: 20 # per page: below this, the page goes to OCR
blank_page_max_ink: 0.0001 # an empty page this clean is blank, not degraded
isolate_worker: true # contain parser/OCR hangs in a child process
timeout_seconds: 120 # whole-document extraction budget
worker_memory_mb: 4096 # address-space cap for the isolated worker
dlp:
quarantine_on_degraded: true # fail closed
hmac_key_env: DLPDUCK_HMAC_KEY
rules:
- include: builtin:default.yaml
- id: local.badge_number # inline rules work too
name: Site badge number
pattern: '\bBADGE-\d{6}\b'
severity: MEDIUM
action: flag
destination:
archive: /srv/dlpduck/archive
quarantine: /srv/dlpduck/quarantine # a SEPARATE mount/ACL in production
work_dir: /srv/dlpduck/work # index/, content/, failed/, spool/
audit:
integrity: chained # chained | none
path: /srv/dlpduck/audit # optional; defaults to work_dir/audit
retention: # opt-in; unset means keep forever
documents_days: null # PDFs, failed queue + the content store
index_days: null
audit_days: null
console:
bind: 127.0.0.1:8080 # IPv6 literals in brackets: "[::1]:8080"
session_secret_env: DLPDUCK_SESSION_SECRET
session_max_age_seconds: 28800 # 8h; sessions can be revoked sooner
session_cookie_secure: false # set true behind TLS
audit_search_terms: hashed # hashed avoids storing search queries
forwarded_allow_ips: null # trusted reverse proxies, e.g. "10.0.0.0/8"
auth:
max_failed_logins: 10 # then that username/address waits out the lockout
lockout_seconds: 300
users: # local/break-glass accounts
- username: breakglass
password_hash: '$argon2id$...' # dlpduck console hash-password
role: dlp_admin
oidc: # the primary path where you have an IdP
issuer: https://idp.example/realms/dlpduck
client_id: dlpduck-console
client_secret_env: DLPDUCK_OIDC_CLIENT_SECRET
roles_claim: roles
role_map:
corp-dlp-team: dlp_admin
plugins:
- name: syslog
args: {host: siem.example, port: 514, protocol: tcp}
- name: webhook
critical: false
args: {url: https://soc.example/hook, secret_env: DLPDUCK_WEBHOOK_SECRET}metadata_fields is an allowlist — only keys named here survive from a
companion file or the PDF's own Info dictionary; an empty list (the default)
keeps nothing. PDF-derived keys are prefixed pdf_ (pdf_title, pdf_author,
…); where both set a key, the companion wins.
A field name is a dot-path into a nested XML or JSON companion document —
device.id reaches a <device><id> grandchild or a {"device": {"id": ...}}
value the same way. No XPath or JSONPath, just descent through nested
objects; the output stays flat, keyed by the literal dotted name.
source:
metadata_fields: [device_id, department, pdf_title, device.site]DLPDuck ships a conservative baseline ruleset inside the package. Include it by
name — not by path, so it resolves the same whether you cloned the repository or
ran pip install:
dlp:
rules:
- include: builtin:default.yaml # the bundled baseline
- include: site-rules.yaml # your own, relative to this config file
- id: pan.generic # redefine anything you disagree with
name: Payment card number
pattern: '\b(?:\d[ -]?){12,18}\d\b'
validator: luhn
action: flag # baseline ships `quarantine`Later definitions win by id, so you never edit the baseline in place —
everything you don't redefine keeps tracking upstream when you update. Paths
without the builtin: prefix resolve relative to the file doing the including.
rules:
- id: fin.iban
name: IBAN
pattern: '\b[A-Z]{2}\d{2}[A-Z0-9]{11,30}\b'
severity: HIGH # INFO | LOW | MEDIUM | HIGH | CRITICAL
action: quarantine # quarantine | flag | ignore
validator: iban_mod97 # luhn | iban_mod97 | nhs_mod11 | none — checksum, kills false positives
scope: line # line | document
mask_keep: 0 # trailing chars left visible; 0 = fully masked
enabled: true
- id: mark.banner_header
name: Classification banner
pattern: 'OFFICIAL-SENSITIVE|SECRET'
severity: CRITICAL
action: quarantine
scope: line
line_scope: page # position window is per page, not per document
min_line: 0
max_line: 3 # only the first four lines of any page
- id: mark.footer_marking
name: Footer marking
pattern: 'RESTRICTED'
from_end: true # count the window from the END of the page
min_line: 0
max_line: 2
- id: cred.password_assignment
name: Password assignment
pattern: '\b\S{8,}\b'
requires_context: # only a hit if this appears nearby
pattern: '(?i)password|passwd|pwd'
within_lines: 1 # ± lines (default 2)Field reference:
| Field | Meaning |
|---|---|
id |
Unique key. Letters, digits, ., _, -; max 64 chars. |
pattern |
Python regex. No implicit flags — write (?i) if you want one. |
severity |
Ranked; the document's highest hit wins. |
action |
quarantine routes the document; flag records it and archives; ignore records the masked hit but never affects routing. |
scope |
line matches per line; document matches the joined text (for values that wrap). |
line_scope |
With a position window: page (per page) or document. |
min_line / max_line |
Inclusive 0-indexed window. Omit both for "anywhere". A document-scope match is placed by the line it starts on. |
from_end |
Count the window from the end (footers). |
validator |
Checksum applied to each candidate before it counts. |
mask_keep |
Trailing characters left visible. Ignored when the match is short enough that a tail would reveal most of it. |
requires_context |
{pattern, within_lines} — a predicate on surrounding lines. |
enabled |
false keeps a rule in the file but out of the ruleset. |
Test a ruleset against a real corpus before deploying it:
dlpduck test-rules ./corpus --config config.yaml # every rule
dlpduck test-rules ./corpus --config config.yaml --rule-id fin.iban| Store | Contents | Purge behaviour |
|---|---|---|
index/ |
Metadata, disposition, and masked findings | Retained as assessment history |
content/ |
Searchable raw text and structured lines | Deleted by purge |
archive/ and quarantine/ |
Original PDFs | Deleted only by hard purge or retention |
dlpduck purge-content <job_id> # delete extracted content
dlpduck purge-content <job_id> --hard # also delete the original PDF
dlpduck compact-index --config config.yaml --commit
dlpduck reindex --config config.yaml --commitStores use UTC dt=YYYY-MM-DD partitions. Configure retention per store; unset windows keep data indefinitely. compact-index merges completed daily index partitions, while reindex reconstructs missing index rows without restoring content previously recorded as purged.
Audit events live under work_dir/audit/dt=YYYY-MM-DD/events.jsonl. With audit.integrity: chained, each event links to the previous event so edits are detectable.
dlpduck verify-audit --config config.yaml
dlpduck redact-audit --config config.yaml \
--seq 4182 --field query --reason "erasure request 41"Redaction replaces selected event fields with [redacted], preserves the chain, and records who performed it and why. Purge and retention record intent before making irreversible changes.
dlpduck console run --config config.yamlScreens: Overview, Jobs, job detail (findings, receipt history, metadata, reveal, reprocess, purge, audit timeline), Needs attention, Search, Correlate, Rules, Audit, Access. Jobs, Search and Correlate share one set of filters (severity, rule, disposition, source, with/without hits, dates) and export the current result as CSV; every search, correlation and export is audited. Overview counters link to their filtered queue; a worker card shows watcher heartbeat, activity and drop-folder backlog.
Needs attention holds refused, failed and interrupted documents — inspect,
retry (individually or in bulk), or mark resolved with a reason. A crash mid
external-delivery requires acknowledging the first attempt may have succeeded
before retrying. Failed documents follow retention.documents_days.
| Permission | viewer | investigator | dlp_admin | auditor |
|---|---|---|---|---|
jobs.list, jobs.metadata.read, rules.read |
✅ | ✅ | ✅ | ✅ |
dlp.hits.read (masked hits, correlation from a hit, rule filter) |
✅ | ✅ | ✅ | |
jobs.text.read (search, correlating a typed value) |
✅ | ✅ | ||
jobs.pdf.read (archived PDFs) |
✅ | ✅ | ||
jobs.pdf.read.quarantined |
✅ | |||
dlp.reveal (cleartext) |
✅ | |||
quarantine.release, jobs.purge, access.write |
✅ | |||
jobs.failed.manage (the refused/unprocessable queue) |
✅ | |||
jobs.reprocess.preview |
✅ | ✅ | ||
jobs.reprocess.commit |
✅ | |||
audit.read, audit.verify |
✅ | ✅ |
A document counts as quarantined if its disposition says so, a release is
pending, or the file physically sits under the quarantine root — one rule
gates both whether its PDF opens and whether a search snippet shows. dlp_admin
is the superuser role; auditor is its opposite (audit trail and masked hits,
never document text, the PDF, or cleartext).
- OIDC is the primary path. With
console.auth.oidcconfigured,/loginredirects to the IdP; roles come fromroles_claimviarole_map. - Local accounts are the break-glass/air-gapped fallback, at
/login?auth=local. Hash a password withdlpduck console hash-password.
| Command | What it does |
|---|---|
validate-config |
Parse config, compile every regex, check paths. Prints warnings too. Use it in CI. |
run |
Start the watcher daemon. |
scan <pdf> |
Dry-run one document; print lines and what would hit. Writes nothing. |
test-rules <dir> |
Run the ruleset over a corpus and report per-rule hit counts. |
search <query> |
Full-text search: all words required, "phrases", -exclusions; filter by severity, rule, disposition, source, hits, dates. --json for scripts. |
jobs |
List current assessments with the same filters, from the index alone. --json for scripts. |
correlate |
Every document holding one sensitive value — from a hit's --hmac, or --value (prompted, never on the command line). Works after content is purged. |
reprocess |
Re-run the current ruleset over already-ingested jobs. Preview by default; --commit to apply. --mode extract re-runs OCR too. |
release <job_id> |
Carry out a pending de-escalation (quarantine → archive). |
purge-content <job_id> |
Soft purge; --hard also deletes the PDF. |
retention |
Show what's past each retention window. Dry run; --apply to delete. |
verify-audit |
Walk the hash chain and report the first break. |
redact-audit |
Empty named fields from one audit event, recording who and why. |
compact-index |
Merge each past day's assessment files into one. Dry-run by default. |
reindex |
Rebuild the index from the content store, or from the PDFs. Disaster recovery — restores where each document is filed, never re-judges it. |
replay-sink <name> |
Drain a sink's spool after an outage. |
console run |
Start the admin console. |
console hash-password |
Hash a password for console.auth.users. |
A plugin is a small class with one method. There are two phases, and the phase you pick decides what your plugin can do, not just when it runs.
enrich |
emit |
|
|---|---|---|
| Runs | after the DLP scan, before disposition | after commit — the PDF is filed, stores written, audit event appended |
| Can change routing? | yes — mutate ctx and the decision follows |
no, the decision is already recorded |
| Typical use | look up a device/department in a directory, attach context | forward to SIEM, webhook, ticketing |
| Failure impact | a critical failure routes the job to failed/ |
a failure is audited and spooled; the job is already safe |
from dlpduck.plugins.base import Plugin
from dlpduck.types import JobContext
class MyPlugin(Plugin):
phase = "enrich" # or "emit"
name = "my_plugin" # appears in audit events and log lines
def __init__(self, some_option: str, name: str | None = None,
critical: bool = False, spool_root=None):
self.some_option = some_option
self.critical = critical
if name:
self.name = name
def run(self, ctx: JobContext) -> None:
...The loader inspects your __init__ signature and passes name, critical, and
spool_root only if you accept them (by name, or via **kwargs). Everything
under the config entry's args: is passed through as keyword arguments. A
TypeError from your constructor becomes a startup config error, not a runtime
surprise — a misconfigured plugin fails validate-config, not document 4,000.
ctx.job_id # blake2b of the PDF bytes — stable, content-derived
ctx.received_at # UTC datetime
ctx.source_name # from config
ctx.pdf_path # staged PDF (enrich phase); moved by the time emit runs
ctx.pdf_sha256
ctx.metadata # dict — the ALLOWLISTED companion/PDF metadata
ctx.text # DocumentText: .lines, .full_text, .page_count, .degraded
ctx.hits # list[DLPHit] — masked_text and match_hmac, never raw values
ctx.disposition # "pending" during enrich; "archive"/"quarantine"/"failed" at emit
ctx.highest_severity
ctx.audit_fields # dict — YOUR output surface; lands in the index row
ctx.errors # list[str] — appended to on plugin failurectx.audit_fields is the enrich phase's product. Write there rather than
mutating ctx.metadata: metadata records what arrived with the document,
audit_fields records what the system worked out about it, and the console shows
them separately.
- Never write a raw match anywhere.
DLPHitdeliberately doesn't carry the matched value — onlymasked_textand a keyedmatch_hmac. If your sink needs to correlate values across documents, forward the HMAC. A plugin that re-derives cleartext fromctx.textand ships it off-box defeats the entire point of the tool. - Be idempotent. Emit plugins can be replayed from the spool.
- Fail loudly, not silently. Raise. The runner audits the failure as
plugin.failedwith your name and the error, appends toctx.errors, and continues — unless you setcritical: true, which stops the job. - Don't block. Set a timeout on anything doing I/O. The pipeline is single-threaded per document.
- Choose
criticaldeliberately.critical: truemeans "if this doesn't happen, the document must not be treated as processed" — right for a compliance-mandated SIEM feed, wrong for a Slack notification.
Delivery-failure handling is inherited — subclass SpoolingSink and implement
build() and deliver(). On failure, the payload is spooled and re-raised; the
same deliver() is reused later by dlpduck replay-sink, so a replayed event
takes exactly the path a live one would have.
# mycompany/dlpduck_plugins.py
from dlpduck.plugins.sinks import SpoolingSink
from dlpduck.types import JobContext
import json, urllib.request
class TicketSink(SpoolingSink):
phase = "emit"
default_name = "ticket"
def __init__(self, endpoint: str, queue: str = "dlp-review", **kwargs):
super().__init__(**kwargs) # name/critical/spool_root/timeout
self.endpoint = endpoint
self.queue = queue
def build(self, ctx: JobContext) -> dict:
# Runs in-process, so keep it cheap and side-effect free.
return {
"queue": self.queue,
"title": f"{ctx.disposition}: {ctx.job_id[:12]}",
"severity": ctx.highest_severity.value if ctx.highest_severity else None,
"rules": sorted({h.rule_id for h in ctx.hits}),
"masked": [h.masked_text for h in ctx.hits], # never h. raw anything
"department": ctx.audit_fields.get("department"),
}
def deliver(self, payload: dict) -> None:
# Raise on failure — the base class spools and re-raises for you.
req = urllib.request.Request(
self.endpoint,
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
method="POST",
)
with urllib.request.urlopen(req, timeout=self.timeout) as resp:
if resp.status >= 300:
raise RuntimeError(f"ticket API returned HTTP {resp.status}")Only quarantined documents raise a ticket? Filter in run():
def run(self, ctx: JobContext) -> None:
if ctx.disposition != "quarantine":
return
super().run(ctx)This one is a deliberate exception to rule #1 above — it ships the document's
full text off-box, to a third-party API, on purpose. That's why it's gated
hard: only archive-disposition documents (never quarantined ones, which are
exactly the sensitive content this tool exists to keep in-house), and only
against an LLM endpoint you've actually reviewed for this — a self-hosted
model, or a vendor under contract — never a default. Think about that gate
before adapting this for your own use.
# mycompany/dlpduck_plugins.py
import json
import os
import urllib.request
from dlpduck.plugins.sinks import SpoolingSink
from dlpduck.types import JobContext
class LlmSummarySink(SpoolingSink):
"""Summarises an archived document's text with an LLM provider and posts
the summary to an internal endpoint. Targets an OpenAI-compatible chat
completions API; adjust build()/deliver() for another provider's shape.
"""
phase = "emit"
default_name = "llm_summary"
def __init__(
self,
api_url: str,
api_key_env: str,
model: str,
summary_endpoint: str,
max_chars: int = 20_000,
**kwargs,
):
super().__init__(**kwargs) # name/critical/spool_root/timeout
self.api_url = api_url
self.api_key = os.environ[api_key_env] # fail at startup, not mid-run
self.model = model
self.summary_endpoint = summary_endpoint
self.max_chars = max_chars
def run(self, ctx: JobContext) -> None:
if ctx.disposition != "archive":
return # never summarise a quarantined document
super().run(ctx)
def build(self, ctx: JobContext) -> dict:
return {"job_id": ctx.job_id, "text": ctx.text.full_text[: self.max_chars]}
def deliver(self, payload: dict) -> None:
summary = self._summarise(payload["text"])
req = urllib.request.Request(
self.summary_endpoint,
data=json.dumps({"job_id": payload["job_id"], "summary": summary}).encode(),
headers={"Content-Type": "application/json"},
method="POST",
)
with urllib.request.urlopen(req, timeout=self.timeout) as resp:
if resp.status >= 300:
raise RuntimeError(f"summary endpoint returned HTTP {resp.status}")
def _summarise(self, text: str) -> str:
req = urllib.request.Request(
self.api_url,
data=json.dumps({
"model": self.model,
"messages": [
{"role": "system", "content": "Summarise this document in three sentences."},
{"role": "user", "content": text},
],
}).encode(),
headers={
"Content-Type": "application/json",
"Authorization": f"Bearer {self.api_key}",
},
method="POST",
)
with urllib.request.urlopen(req, timeout=self.timeout) as resp:
body = json.loads(resp.read())
return body["choices"][0]["message"]["content"]plugins:
- path: mycompany.dlpduck_plugins.LlmSummarySink
args:
api_url: https://api.your-llm-provider.example/v1/chat/completions
api_key_env: DLPDUCK_LLM_API_KEY
model: your-provider-model-id
summary_endpoint: https://intranet.example/dlp-summaries
timeout: 30class LdapEnrich(Plugin):
phase = "enrich"
name = "ldap_enrich"
def __init__(self, server: str, key_field: str = "device_id",
name: str | None = None, critical: bool = False, spool_root=None):
self.server = server
self.key_field = key_field
self.critical = critical
if name:
self.name = name
def run(self, ctx: JobContext) -> None:
key = ctx.metadata.get(self.key_field)
if key is None:
return # nothing to look up isn't an error
ctx.audit_fields.update(self._lookup(key))Because this runs before disposition, whatever it writes to ctx.audit_fields is
visible to everything downstream and is stored on the index row. See
dlpduck/plugins/enrich.py for a working static-mapping version.
Built-ins are selected by name; anything else needs a dotted path that's
importable from the running environment:
plugins:
- name: syslog # built-in
args: {host: siem.example, protocol: tcp}
- path: mycompany.dlpduck_plugins.TicketSink
critical: false
args:
endpoint: https://tickets.example/api/v1/issues
queue: dlp-review
- path: mycompany.dlpduck_plugins.LdapEnrich
enabled: false # keep the config, skip the plugin
args: {server: ldaps://dc1.example}Built-in names: syslog, webhook, static_enrich.
webhook posts to exactly the URL you configure and refuses to follow
redirects — a redirect target returning 200 would otherwise read as a
successful delivery while nothing was actually received, and the signing
header would travel to a host you didn't choose. A redirect spools as a
delivery failure instead, retryable with replay-sink.
Then verify before you deploy:
dlpduck validate-config --config config.yaml # loads and constructs every pluginPlugins take a plain JobContext, so they test without a pipeline:
def test_ticket_sink_omits_raw_values():
ctx = make_context(hits=[hit(rule_id="pan.generic", masked_text="••••1111")])
payload = TicketSink(endpoint="http://x").build(ctx)
assert "4111" not in json.dumps(payload)Apache License 2.0 — see LICENSE.
The console ships the IBM Plex typeface (dlpduck/console/static/fonts/), which
is licensed separately under the SIL Open Font License 1.1 — see
fonts/LICENSE.txt. It is bundled
rather than fetched from a CDN so the console works with no route to the
internet, and emits no telemetry to a third party.
