Compare commits
49 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 2dee3b722c | |||
| 8daa7a1991 | |||
| f89ca2aceb | |||
| 57a6fb5daa | |||
| 49dafedbbc | |||
| b83523fc92 | |||
| 8ae5c332c3 | |||
| 420f9eb9a0 | |||
| 4314b11237 | |||
| eaea520020 | |||
| b8d22904f1 | |||
| cbd16ce3f6 | |||
| 017cfd7e28 | |||
| 954173cacf | |||
| 97934d641a | |||
| afbe6367cc | |||
| af38583b53 | |||
| eeb72fd677 | |||
| 320cb48b5d | |||
| f2a61ad88c | |||
| 8d84f9fe8c | |||
| e38e3c563c | |||
| a30edff248 | |||
| be3820bbbd | |||
| 51ce0bf9eb | |||
| 6ec900e456 | |||
| d4304869b6 | |||
| 9aae3097e7 | |||
| 2ff2074755 | |||
| 1698cf8d1e | |||
| b94a74ef15 | |||
| 9a8a3c696f | |||
| bedf64e645 | |||
| 586d2c07a1 | |||
| 2ba83c3aed | |||
| 1352f0d54d | |||
| a958f7cb37 | |||
| 457a58dcc5 | |||
| 2af57065c6 | |||
| 564b011c7a | |||
| 9eb7aadced | |||
| c6bceaef37 | |||
| dc63d6acaf | |||
| 8937f35f00 | |||
| 0990f2a90d | |||
| 0cc002cdfa | |||
| 2c9e899762 | |||
| 2e416f96cf | |||
| 9269aa99f7 |
83
.agents/skills/security-audit/AI-AND-LLM.md
Normal file
83
.agents/skills/security-audit/AI-AND-LLM.md
Normal file
@@ -0,0 +1,83 @@
|
||||
# AI, LLM, and Agent Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when a language model participates in a trust-sensitive decision: chatbots and assistants, RAG pipelines, persistent agent memory, agent/tool-calling loops, MCP servers and clients, code that builds prompts from untrusted input, or code that consumes model output and acts on it. The important data flow is *untrusted content → model or memory → capability, authority, or sink*.
|
||||
|
||||
Use this alongside `ATTACK-CLASSES.md`, not instead of it. Transport, access control, query construction, filesystem use, and output rendering remain ordinary trust boundaries. This file covers the model-specific delegation layer. Split large targets by retrieval, memory, tool dispatch, MCP, and output handling.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- Prompt injection alone is not a finding. Require a code-level boundary failure: content reaches another principal's context, invokes authority the requester lacks, discloses data they cannot read, or drives a sink they cannot reach directly.
|
||||
- Model output, memory, tool descriptions, and MCP responses are untrusted inputs. Point to the code that grants authority, trusts output, writes durable state, or feeds a sink.
|
||||
- A guardrail prompt is not a security boundary. Count only deterministic checks, resource-scoped authorization, isolation, binding, and constrained credentials.
|
||||
- State the attacker, affected principal, effective execution identity, resource, exact action, authority used, and observable impact. An intentional direct request to use the requester's existing authority is not a delegation defect merely because a model executes it.
|
||||
- Authorization and action binding are separate controls. Attacker-controlled content that causes an action under an affected principal's valid authority is an action-binding failure when that principal did not intentionally request or approve the exact action.
|
||||
- Classify every candidate as `confirmed` only after source evidence and bounded local validation establish the boundary and result. Use `needs_validation` when a required provider, deployment, model, renderer, or identity behavior is not observable locally.
|
||||
```
|
||||
|
||||
## Context, retrieval, and memory attack classes (subagent_type: `general`)
|
||||
|
||||
**Indirect injection through retrieved or ingested content**
|
||||
An attacker can write a RAG document, indexed page, file, email, issue body, tool response, or metadata that enters a different principal's model context. Trace who can write each source, how retrieval scopes it, whose session consumes it, and what capability is enabled there. Check isolation, resource authorization, and binding to the consuming principal's intent separately. The defect is a missing deterministic control, not persuasive text by itself.
|
||||
|
||||
**Cross-session or cross-tenant context bleed**
|
||||
Conversation history, embeddings, retrieval results, or prompt caches are keyed too broadly. Verify tenant and ACL filters in the query itself and every cache key. A tenant field stored on an object is not enforcement if an alternate query, shared cache, or batch path omits it.
|
||||
|
||||
**Persistent memory poisoning**
|
||||
Attacker-controlled content or model summaries are written into memory that later shapes another task, user, or privileged session. Review who may create, update, merge, and delete memory; its provenance and tenant scope; whether low-trust observations become durable instructions or facts; and whether retrieval distinguishes user preferences from tool policy. Memory intentionally saved by a user and used only for that user's intentional, allowed requests is not a cross-boundary finding.
|
||||
|
||||
**Prompt role and provenance confusion**
|
||||
Prompt assembly lets untrusted text impersonate a system message, prior turn, tool result, policy, or memory record. Look for string concatenation, untyped history, caller-controlled role fields, and serialization round trips that lose source labels. Confirm that the forged provenance changes a deterministic trust decision or reaches a meaningful capability.
|
||||
|
||||
## Tool and action attack classes (subagent_type: `general`)
|
||||
|
||||
**Tool-argument injection into a downstream sink**
|
||||
Model-produced arguments reach SQL, shell, file, URL-fetch, or privileged APIs without handler-side validation. Treat the tool schema as input parsing, then follow each field from decoded call to sink. Structured output narrows shape; it does not establish authorization, safe paths, safe URLs, or query semantics.
|
||||
|
||||
**Excessive agency and confused-deputy authority**
|
||||
The agent uses a service identity or broad credential, while the tool handler does not re-check the requesting principal's permission on the named resource. Verify both the effective identity and whether the caller could perform that exact operation through the normal product interface. A shared credential with enforced per-user query scope is not a defect.
|
||||
|
||||
**Action-confirmation and approval binding**
|
||||
A user approves one described action but execution can use changed arguments, a different resource, a different principal, or a later model turn. An action-binding defect also exists when attacker-controlled content causes a side effect under a victim's valid authority without the victim's intentional request or approval, even if generic authorization permits the victim to perform it. Review whether intent or confirmation binds the normalized tool name, complete argument object, requester, target, amount, expiry, and batch membership. Check retries and resumed sessions: an approval must not authorize a mutated or duplicate side effect.
|
||||
|
||||
**Tool-schema and dispatcher disagreement**
|
||||
The schema accepts aliases, extra fields, duplicate keys, coercions, nested free-form objects, or out-of-range values that the dispatcher or handler interprets differently. Compare schema validation, canonicalization, generated bindings, and handler defaults. Validate again where values become resource selectors or security-relevant options.
|
||||
|
||||
**Unbounded delegated action loops**
|
||||
A bounded request can enqueue repeated spend, send, mutation, or external API work without a per-request budget, per-action authorization, cancellation, or idempotency control. Confirm impact on shared cost, quotas, other users, or durable state. Do not test by exhausting a service; use code-level accounting and a locally bounded loop.
|
||||
|
||||
## MCP and sub-agent trust classes (subagent_type: `general`)
|
||||
|
||||
**Sub-agent and MCP trust inheritance**
|
||||
A delegated task receives the full session, credentials, memory, or capabilities rather than the least authority required. Check the principal and tenant carried into each call, capability narrowing, credential audience, and whether delegated results are treated as untrusted on return.
|
||||
|
||||
**MCP server and tool identity confusion**
|
||||
Calls or results are routed by attacker-influenceable server names, tool names, request IDs, resource URIs, or model-selected aliases rather than the authenticated connection and outstanding request. Check whether two servers can claim the same tool or resource identity, whether reconnect changes the binding, and whether a response from one server can satisfy another server's pending call.
|
||||
|
||||
**MCP metadata and schema as policy**
|
||||
Tool descriptions, resource metadata, prompts, completion hints, or schemas supplied by an MCP peer are trusted as policy or authorization. These fields can guide the model but cannot grant capability. Find the deterministic allowlist, server identity check, and handler authorization that remain authoritative when metadata conflicts.
|
||||
|
||||
## Output and disclosure attack classes (subagent_type: `general`)
|
||||
|
||||
**Insecure output rendering**
|
||||
Model output reaches an executing HTML, Markdown, template, URL, or command sink without the sink's required encoding and policy. For browser rendering, verify auto-loaded resources and CSP or sanitization in `CLIENT-SIDE.md`; renderer behavior outside the repository makes the candidate `needs_validation`.
|
||||
|
||||
**Sensitive context extraction**
|
||||
The assembled context contains credentials, another user's data, private source, or policy values that themselves grant access, and user-influenced output exposes them. Read prompt assembly and data-fetch code. Disclosure of generic instructions or behavior that does not cross a data boundary is not a finding.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Draw four maps first: each execution identity, each capability, every writable context or memory source, and each output destination. Then connect the principal at the start to the authority at the end.
|
||||
- Start at side-effecting tools and work backward through dispatcher, schema, confirmation, model context, retrieval, and ingestion. Start at durable memory reads and trace every writer.
|
||||
- Compare direct, queued, retry, resume, batch, and delegated paths for the same action. The strongest gate must apply after arguments are final and before every side effect.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Name the crossed boundary and observable result: attacker, affected principal or shared resource, execution identity, target, and unauthorized or unrequested action or disclosure.
|
||||
2. For confused-deputy authority claims, prove the tool lacks requester-and-resource authorization and that the attacker cannot perform the same action normally. For action-binding claims, instead prove attacker-controlled content caused an action under the affected principal's authority that the principal did not intentionally request or approve. Valid generic authorization does not establish that intent.
|
||||
3. For memory or retrieval claims, cite both the attacker-controlled write and the later cross-principal read or privileged decision. A shared record without a reachable consumer is not enough.
|
||||
4. For action binding, establish the intentional request or normalized approved object, if any, and compare it with the object the handler uses. Confirm a locally observable unrequested action, mutation, duplicate, or authority change without extending the test into harmful execution. For schema disagreement, compare the normalized validated object with the handler's object.
|
||||
5. For MCP identity claims, verify the authenticated connection, request correlation, tool namespace, and effective credential. Mark `needs_validation` if external server identity or deployment routing is required.
|
||||
6. Return `confirmed` findings only with a complete source trace and meaningful result. Return `needs_validation` for a specific unresolved boundary fact and state the bounded local or owner-observed check needed to resolve it.
|
||||
130
.agents/skills/security-audit/ATTACK-CLASSES.md
Normal file
130
.agents/skills/security-audit/ATTACK-CLASSES.md
Normal file
@@ -0,0 +1,130 @@
|
||||
# Attack Classes
|
||||
|
||||
#### Attack classes — choose and split based on Phase 1
|
||||
|
||||
Select attack classes relevant to the application type. Not every class applies to every codebase. The list below is a starting point; add application-specific classes from Phase 1 and split large codebases per subsystem. Frame work as finding, validating, fixing, and prioritizing vulnerabilities. Keep validation to source review and bounded local fixtures; do not develop payload chains, test availability on live services, or take action in shared environments.
|
||||
|
||||
Use `confirmed` only when source evidence and bounded validation establish the full boundary and meaningful result. Use `needs_validation` when a specific deployment, provider, platform, identity, or runtime fact is unavailable; state the missing fact and the safe owner-observed or local check that resolves it.
|
||||
|
||||
> **Native / binary / kernel targets** (C/C++/Rust-unsafe, kernel modules, parsers and decoders, FFI, concurrent runtimes, binary loaders, JITs, firmware): use the memory-safety, integer/ABI, concurrency, binary-loader, and privileged-interface classes in [MEMORY-SAFETY-AND-BINARY.md](MEMORY-SAFETY-AND-BINARY.md).
|
||||
>
|
||||
> **AI / LLM / agent targets** (chatbots, RAG, persistent memory, tool-calling agents, MCP servers/clients, prompt assembly, or model-controlled actions): use the context, memory-poisoning, action-binding, tool-schema, MCP-identity, and output classes in [AI-AND-LLM.md](AI-AND-LLM.md).
|
||||
>
|
||||
> **HTTP, web, and identity targets** (ordinary web apps, APIs, reverse proxies, CDNs, gateways, custom HTTP parsers, sessions, CSRF, JWT, OAuth/OIDC, SAML, MFA, passkeys, account recovery/linking, API keys, or mTLS): use [WEB-PROTOCOL-AND-AUTH.md](WEB-PROTOCOL-AND-AUTH.md).
|
||||
>
|
||||
> **Client-side and browser targets** (SPAs, browser extensions, embedded webviews, service workers, browser storage, cross-window messaging, CORS, WebSockets, or DOM rendering): use [CLIENT-SIDE.md](CLIENT-SIDE.md).
|
||||
>
|
||||
> **Supply-chain and release targets** (dependency resolution, generated inputs, CI, release/signing/promotion, updates, plugins, or extensions): use [SUPPLY-CHAIN-AND-RELEASE.md](SUPPLY-CHAIN-AND-RELEASE.md).
|
||||
>
|
||||
> **Cloud and deployment targets** (IAM, infrastructure as code, containers/Kubernetes, service mesh, serverless/edge, ingress, provider events, or runtime configuration): use [CLOUD-AND-DEPLOYMENT.md](CLOUD-AND-DEPLOYMENT.md).
|
||||
>
|
||||
> **Protocol, RPC, and messaging targets** (gRPC, GraphQL transports, Protobuf/Cap'n Proto/Thrift, custom protocols, queues, brokers, pub/sub, webhooks, or streaming RPC): use [PROTOCOLS-RPC-AND-MESSAGING.md](PROTOCOLS-RPC-AND-MESSAGING.md).
|
||||
>
|
||||
> **Resource-exhaustion and availability targets** (untrusted work can consume shared CPU, memory, disk, connections, workers, queues, quotas, or operator-owned spend): use [RESOURCE-EXHAUSTION-AND-AVAILABILITY.md](RESOURCE-EXHAUSTION-AND-AVAILABILITY.md).
|
||||
>
|
||||
> **Data-isolation and lifecycle targets** (multi-tenant stores, caches/search, object links, analytics, export/backup, migration, deletion, retention, or restore): use [DATA-ISOLATION-AND-LIFECYCLE.md](DATA-ISOLATION-AND-LIFECYCLE.md).
|
||||
>
|
||||
> **Desktop, mobile, and local-IPC targets** (native apps, deep links, webview bridges, exported components, privileged helpers, local daemons, Unix sockets/XPC/Binder/D-Bus): use [DESKTOP-MOBILE-AND-LOCAL-IPC.md](DESKTOP-MOBILE-AND-LOCAL-IPC.md).
|
||||
|
||||
**Injection** (subagent_type: `general`)
|
||||
Trace untrusted input from entry point to dangerous sink. What counts as a "dangerous sink" depends on the application:
|
||||
- Web apps: SQL queries, HTML output, shell commands, template engines, file paths, HTTP redirects, deserialization
|
||||
- Libraries: any function that processes caller-supplied data without validation — buffer operations, parsers, format strings
|
||||
- CLI tools: shell command construction, file path handling, environment variable interpolation
|
||||
- Services: query construction, message serialization, log injection, LDAP/XPATH queries
|
||||
- Client-side (browser/JS): DOM XSS, prototype pollution, `postMessage`/origin trust, and other browser-side classes — covered by the [CLIENT-SIDE.md](CLIENT-SIDE.md) companion blocks when selected
|
||||
|
||||
Do not stop at the obvious direct paths. Look for indirect injection: data stored safely, then retrieved and used in a dangerous context by different code. Look for injection through field names, keys, headers, and metadata — not just values. Look for injection into secondary systems (logs, caches, search indexes, analytics).
|
||||
|
||||
**Access control** (subagent_type: `general`)
|
||||
Verify that a caller cannot do something outside its authority. Go beyond checking whether permission checks exist — verify they check the *right* permission for the *right* resource via the *right* mechanism:
|
||||
- Is there a path to the same state change that checks a different (weaker) permission?
|
||||
- Can a field in the request body override what the permission system intended to restrict?
|
||||
- Are there endpoints that gate on authentication but forget authorization?
|
||||
- Does the same resource have multiple access paths with inconsistent checks?
|
||||
- What about bulk/batch/export/import operations — do they enforce per-item permissions?
|
||||
|
||||
For complex access models, split into separate agents for auth bypass vs authorization logic.
|
||||
|
||||
**Resource and file handling** (subagent_type: `general`)
|
||||
- Path traversal (reading/writing outside intended directories) — including through symlinks, encoded sequences, and null bytes
|
||||
- SSRF (making the application fetch attacker-controlled URLs) — including through redirects, DNS rebinding, and URL parser differentials
|
||||
- Unsafe deserialization, archive extraction (zip slip), temp file handling
|
||||
- Memory safety (if applicable): buffer overflows, use-after-free, integer overflow
|
||||
- Race conditions on file operations (TOCTOU between check and use)
|
||||
|
||||
**Cryptography and secrets** (subagent_type: `general`)
|
||||
- Weak randomness for security-critical values (tokens, keys, nonces)
|
||||
- Hardcoded secrets, secrets in logs, error messages, URLs, or client-visible responses
|
||||
- Broken key derivation, missing HMAC verification, nonce reuse
|
||||
- Timing side-channels on secret comparison
|
||||
- Misuse of crypto primitives (ECB mode, unauthenticated encryption, static IVs, etc.)
|
||||
- What happens when crypto operations fail? Does the error path fall back to no-crypto?
|
||||
|
||||
**Business logic** (subagent_type: `general`)
|
||||
Hunt logic errors by hand: standard scanners cannot find them, and they yield high-impact findings. For each major workflow:
|
||||
- **State machine violations**: Can you skip steps? Go backwards? Reach an invalid state? What happens if you replay a completed flow? What about partial failure — if step 2 of 3 fails, is step 1 rolled back?
|
||||
- **Race conditions with business impact**: Concurrent operations that produce invalid states (double-spend, double-approve, lost updates). Focus on operations that check-then-act non-atomically.
|
||||
- **Numeric/quantity manipulation**: Negative values, zero values, overflow, precision loss, type coercion between string and number.
|
||||
- **Access boundary violations**: Not "does the permission check exist" but "is it the right check for the business rule?" Can input to one operation bypass a restriction enforced on a different operation for the same effect?
|
||||
- **Implicit trust assumptions**: Data from storage, config, other components, or plugins assumed safe because "we validated it on the way in." What if a different code path wrote it?
|
||||
- **Time-based logic**: Expiry checks, scheduling, rate windows, clock skew. What happens at exact boundary moments? What about timezone differences between components?
|
||||
- **Default and fallback behavior**: What is the security posture when config is missing? When a feature flag is off? When a dependency is unavailable? When the system is mid-migration?
|
||||
|
||||
**Feature abuse and data leakage** (subagent_type: `general`)
|
||||
Legitimate features used for unintended purposes. Look for bugs in the design, not only in the code:
|
||||
- **Export/backup as exfiltration**: Can a low-privilege user trigger an export, snapshot, or backup that includes data above their access level? Can they export other users' data? Does the export include deleted/draft/private content? Revision history that was supposed to be pruned?
|
||||
- **Import/restore as injection**: Can import overwrite existing data? Can it create records that bypass normal validation? Can it inject content into collections the user has no write access to? Does it respect the same permission model as the UI?
|
||||
- **Search/filter/sort as oracle**: Can search queries reveal whether content exists that the user cannot directly access? Do filter parameters let users probe statuses, roles, or fields they should not know about? Does sorting by a hidden field reveal its values through result ordering?
|
||||
- **Enumeration through side effects**: Do error messages differ between "does not exist" and "no access"? Do response times differ? Response sizes? HTTP status codes? Can you enumerate users through password reset, invite, or registration flows?
|
||||
- **Preview/draft/staging leakage**: Are preview tokens scoped to one item or do they unlock broader access? Can draft content be discovered through search, RSS feeds, sitemaps, or API listing endpoints? Can cache headers cause a CDN to serve private content publicly?
|
||||
- **Notification/webhook as SSRF**: Can a user set a notification URL, webhook URL, or callback URL that the server fetches? Is it validated against internal networks? What about after a redirect?
|
||||
|
||||
**Chained vulnerabilities and trust boundaries** (subagent_type: `general`)
|
||||
Individually allowed or contained behavior can become a vulnerability when another component or lifecycle step relies on a stronger guarantee:
|
||||
- **Multi-step boundary failures**: Map what a low-privilege principal may read, write, invoke, and retain, then connect only concrete outputs to later trust decisions. Confirm each prerequisite and do not assume a downstream effect.
|
||||
- **Cross-component trust gaps**: Component A validates input and passes it to component B. Compare the exact guarantee A produces with what B assumes, including truncation, type coercion, normalization, tenant scope, and plugin/extension access.
|
||||
- **Second-order use**: Data safe when stored may become dangerous in a later context. A field name becomes a JSON path, a slug becomes a file path, escaped text enters raw rendering, or a stored string becomes a URL, regex, template, or policy expression.
|
||||
- **Scope and capability growth**: Token, API-key, plugin, OAuth, MCP, or AI capabilities become broader after delegation, refresh, caching, role change, or composition. Name the concrete operation the resulting principal should not have.
|
||||
- **Timing and ordering**: Review setup, migration, soft-delete, revoke/cache expiry, check/use, and validate/consume windows. Confirm stale state is accepted before reporting.
|
||||
- **Rollback and recovery**: Undelete, restore, revision rollback, and cancellation must apply current ownership, validation, and authorization. Confirm which invalid state is restored.
|
||||
|
||||
**Wildcard** (subagent_type: `general`)
|
||||
You are not given a category. Find vulnerabilities outside the standard classes already assigned.
|
||||
|
||||
Read code that looks boring or disconnected from security. Follow incomplete, experimental, compatibility, and fallback features, but retain the same concrete boundary and validation requirements as every other class.
|
||||
|
||||
Use these starting points, but do not limit yourself to them:
|
||||
- What is the strangest code in the codebase? Why does it exist? What happens if it is abused?
|
||||
- Are there any features that feel half-finished, experimental, or bolted on? Those have the weakest security because they got the least review.
|
||||
- What happens if you use the API in a way the frontend never would? The UI constrains users, but the API does not. What API calls are possible but never made by the client?
|
||||
- Are there any hidden or undocumented endpoints, parameters, headers, or features? Look at route registrations, middleware, and config for things that are not in the docs.
|
||||
- What happens when you mix features that were not designed to work together? Localization + preview + caching. Import + plugins + webhooks. OAuth + impersonation + API keys.
|
||||
- Is there anything interesting in the git history? Reverted security fixes, commented-out auth checks, secrets that were committed then removed (still in history).
|
||||
- Which valid-account actions affect other users, shared integrity, availability, or operator-owned cost? Verify containment, quotas, authorization, and recovery around those actions.
|
||||
- Which operations are irreversible or require elevated confirmation? Bind authorization and approval to the final principal, action, and resource.
|
||||
- What assumptions does the code make about the environment? That the database is local, that the clock is accurate, that DNS is trustworthy, that the filesystem is case-sensitive?
|
||||
- Look at the test files — what are they **not** testing? Compare the edge cases the developer thought about (tests exist) with the ones they did not (no tests).
|
||||
|
||||
Pursue anomalies inside your assigned scope until the invariant is settled. If something looks strange, read it until you can state whether it is safe. If a function has a comment explaining why it is safe, verify the explanation. If a variable is named `temp` or `hack` or `legacy`, read it closely.
|
||||
|
||||
**Obvious things** (subagent_type: `general`)
|
||||
Other agents hunt subtle bugs. This agent checks the basic exposures that are easy to overlook because everyone assumes someone else already checked them:
|
||||
- Are there any hardcoded passwords, API keys, tokens, or secrets in the source? (grep for `password`, `secret`, `apikey`, `token`, `Bearer`, `-----BEGIN`, common default passwords)
|
||||
- Are there any TODO/FIXME/HACK/XXX comments that reference security? (`TODO: add auth`, `FIXME: validate input`, `HACK: skip permission check`)
|
||||
- Is debug mode / dev mode properly gated? Can it be enabled in production via environment variable, query parameter, or header?
|
||||
- Are there test/example/seed credentials that work in production?
|
||||
- Is there a `/debug`, `/admin`, `/test`, `/status`, `/health`, `/metrics`, `/env`, `/.env`, `/config` endpoint that is unprotected?
|
||||
- Are there any `.env`, `.env.local`, `credentials.json`, `*.pem`, `*.key` files checked into the repo?
|
||||
- Does the `.gitignore` actually cover secrets, uploads, and local config?
|
||||
- Are dependencies pinned? Are there known CVEs in the dependency tree? (check lockfiles)
|
||||
- Are there any `eval()`, `exec()`, `child_process`, `Function()`, `vm.runInContext`, `import()` with dynamic input?
|
||||
- Are CORS headers set to `*` or overly permissive? Is `Access-Control-Allow-Credentials` combined with a wildcard origin?
|
||||
- Are cookies missing `HttpOnly`, `Secure`, or `SameSite` attributes?
|
||||
- Are there any open redirects? (parameters named `redirect`, `return`, `next`, `url`, `goto`, `continue` that feed into redirects without validation)
|
||||
- Is TLS enforced? Are there any HTTP-only endpoints?
|
||||
- Are error responses in production returning stack traces, internal paths, or SQL errors?
|
||||
|
||||
This agent does not need to be creative. It needs to be thorough and literal. Check every item. Report each result.
|
||||
|
||||
**Important**: For any finding this agent reports, it must verify the full code path, not just surface appearance. If a cookie is missing `HttpOnly`, check whether the cookie contains security-sensitive data and whether JS needs to read it by design. If an error message contains a field name, check whether the field is ever actually populated with sensitive data. A flag is not a finding — trace the impact before reporting.
|
||||
83
.agents/skills/security-audit/CLIENT-SIDE.md
Normal file
83
.agents/skills/security-audit/CLIENT-SIDE.md
Normal file
@@ -0,0 +1,83 @@
|
||||
# Client-Side and Browser Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when meaningful trust decisions or untrusted rendering happen in a browser: single-page apps, browser extensions, embedded webviews, service workers, offline applications, and code that renders attacker-influenceable content into the DOM, receives cross-window messages, or uses browser storage. These paths include sources the server never sees, such as URL fragments, `window.name`, `postMessage`, and previously cached content.
|
||||
|
||||
Use alongside `ATTACK-CLASSES.md`. This file covers browser sources and sinks, origin boundaries, browser persistence, and cross-site state oracles. Use `DESKTOP-MOBILE-AND-LOCAL-IPC.md` for the native side of a webview bridge, and `WEB-PROTOCOL-AND-AUTH.md` for server-side CSRF, sessions, and auth callbacks.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- A client-side candidate needs a controllable source and an executing or disclosing sink. Name both and show attacker-influenced data reaching the sink.
|
||||
- The impact must reach a victim's session, another origin, or shared persistence. Self-injection and disclosure of the attacker's own data are not findings.
|
||||
- Framework escaping, browser same-origin policy, CSP, COOP/CORP, service-worker scope, and modern noopener defaults are real controls. Verify them before assigning impact.
|
||||
- Browser storage and caches are shared by origin and may outlive login state. Identify who writes, who reads, and which account, tenant, or worker lifecycle clears each record.
|
||||
- Use `confirmed` only for complete source evidence plus bounded local browser tests. Use `needs_validation` when renderer, extension permission, deployed header, or browser-policy behavior is required but unavailable.
|
||||
```
|
||||
|
||||
## DOM and object-state attack classes (subagent_type: `general`)
|
||||
|
||||
**DOM-based XSS**
|
||||
Trace `location` fields, `document.referrer`, `window.name`, message data, storage, and browser-controlled document state into `innerHTML`, `outerHTML`, `document.write`, string-evaluating APIs, executable URLs, jQuery HTML APIs, or framework escape hatches. Interpolation escaped by the framework is not a finding.
|
||||
|
||||
**DOM clobbering**
|
||||
Attacker-injected `id` or `name` attributes shadow a global, form property, configuration object, or initialization flag later trusted by code. Require both a markup path that preserves the attribute and a security-relevant use of the clobbered value.
|
||||
|
||||
**Prototype pollution and gadget chain**
|
||||
An attacker-controlled key reaches a recursive write such as deep merge or path assignment and modifies prototype state. Then a reachable gadget consumes the polluted property to change authorization, execution, navigation, or rendering. `JSON.parse`, a shallow copy, or pollution without a gadget is not enough.
|
||||
|
||||
## Cross-origin messaging and network attack classes (subagent_type: `general`)
|
||||
|
||||
**`postMessage` origin and source trust**
|
||||
A handler performs a sensitive action with `event.data` without an exact origin allowlist and, where multiple frames share an origin, the expected `event.source`. On the send side, sensitive data sent to `*` reaches an unintended embedder. Weak substring, prefix, suffix, or unanchored-regex origin matching is not an origin check.
|
||||
|
||||
**Cross-site WebSocket request use**
|
||||
A WebSocket upgrade accepts ambient cookies from an untrusted origin without an `Origin` check or channel-specific token, allowing the victim's session to read or mutate data. Confirm both the upgrade behavior and a security-relevant message handler.
|
||||
|
||||
**Credentialed CORS trust**
|
||||
The server reflects or weakly matches `Origin` while allowing credentials and returns sensitive responses. A bare wildcard with credentials is rejected by browsers; report only the actual reflected/allowed origin path and cross-origin data or mutation.
|
||||
|
||||
## Service-worker and browser-storage attack classes (subagent_type: `general`)
|
||||
|
||||
**Service-worker registration and scope takeover**
|
||||
Attacker-influenceable content can become the registered worker script, control a path that receives an over-broad `Service-Worker-Allowed` scope, or alter update imports without integrity control. Verify the final script URL, response MIME type, origin, scope, and who controls every imported script. A normal same-origin worker with intended scope is not a defect.
|
||||
|
||||
**Service-worker cache and identity confusion**
|
||||
The worker caches personalized responses without including account, tenant, authorization state, or request mode in its policy, then serves them after account switch or logout. Review fetch-event routing, cache names and keys, navigation fallbacks, cache cleanup, and whether error/offline paths return another user's prior response.
|
||||
|
||||
**Browser-storage disclosure and stale authorization**
|
||||
Tokens, private responses, draft data, or authorization decisions remain in `localStorage`, `sessionStorage`, IndexedDB, Cache Storage, extension storage, or client state and become readable by another account or less-trusted same-origin component. Storage of a token alone is not a finding; require a realistic reader with less authority, or continued use after revocation/logout.
|
||||
|
||||
**Cross-context storage and broadcast confusion**
|
||||
`storage` events, `BroadcastChannel`, shared workers, or origin-wide caches carry identity or commands between tabs without binding them to the current session. Check account switching, private/public windows, tenant changes, and stale tabs that can overwrite newer auth state.
|
||||
|
||||
## Cross-site information leak classes (subagent_type: `general`)
|
||||
|
||||
**XS-Leaks and cross-origin state oracles**
|
||||
An attacker page can distinguish protected cross-origin state through resource load/error events, frame or window state, redirect behavior, timing, cache state, or response size while the browser attaches victim credentials. Require one concrete secret-bearing predicate such as whether a private object, role, or account exists. Generic timing variance or public-resource availability is not a finding.
|
||||
|
||||
**Window and opener state disclosure**
|
||||
A cross-origin window's permitted metadata or navigation result reveals protected state, or a retained opener/named-window relationship lets an attacker-controlled page influence a privileged navigation. Check COOP, frame protections, `noopener`, exact origin, and whether the observable state is confidential.
|
||||
|
||||
## UI-redress and navigation attack classes (subagent_type: `general`)
|
||||
|
||||
**Clickjacking**
|
||||
A framed, state-changing action lacks effective `frame-ancestors`, `X-Frame-Options`, or equivalent UI isolation. Require the sensitive action and confirm it can complete in the framed state; missing headers on read-only content are hardening notes.
|
||||
|
||||
**Client-side navigation confusion**
|
||||
A client source controls redirect or navigation without scheme and destination policy, including executable `javascript:` or `data:` destinations. Reverse tabnabbing applies only where code explicitly keeps `window.opener`, uses `window.open` without isolation, or supports a browser without implicit `noopener`.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Start from DOM, navigation, worker, message, and storage sinks, then trace backward to browser-only and server-controlled sources. Record the browser policy that should stop the path.
|
||||
- Test account switch, logout, worker update, offline fallback, and stale-tab state with a local test origin and dummy accounts. Do not use production users, origins, or shared services.
|
||||
- For XS-Leaks, list only predicates proved by source and local browser behavior. Then identify the response headers or rendering choice that would remove the oracle.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Cite the source, sink, browser policy, affected origin/session, and observable mutation or disclosure.
|
||||
2. For prototype pollution, prove the recursive write and a security-relevant gadget. For DOM clobbering, prove the markup survives and the shadowed value is used.
|
||||
3. For service workers and storage, prove lifecycle reachability: an attacker-controlled write or cache entry must reach a different account, tenant, or later authorization state.
|
||||
4. For messaging, CORS, WebSocket, and XS-Leaks, show exact origin/source validation and the protected state or action exposed. Confirm that CSP, COOP/CORP, cookies, and SameSite policy do not already block it.
|
||||
5. Return `confirmed` findings only with a complete client path and bounded local evidence. Return `needs_validation` with the precise deployed header, extension permission, browser version, or renderer behavior an owner must verify.
|
||||
86
.agents/skills/security-audit/CLOUD-AND-DEPLOYMENT.md
Normal file
86
.agents/skills/security-audit/CLOUD-AND-DEPLOYMENT.md
Normal file
@@ -0,0 +1,86 @@
|
||||
# Cloud and Deployment Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when the repository defines cloud identity, infrastructure, containers, Kubernetes, service mesh, serverless functions, edge workers, ingress, object storage, managed services, or environment-specific configuration. This domain asks whether deployed components receive the intended identity, isolation, network reachability, secrets, and policy. Source often expresses intent rather than live fact, so separate source-confirmed defects from deployment validation needs.
|
||||
|
||||
Use `SUPPLY-CHAIN-AND-RELEASE.md` for build and promotion trust, `WEB-PROTOCOL-AND-AUTH.md` for HTTP proxy semantics, and `DATA-ISOLATION-AND-LIFECYCLE.md` for data-store tenant scope.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- Do not infer a live exposure from a manifest alone. Establish which environment consumes it, what defaults or overlays modify it, and whether the source path is active.
|
||||
- Map each workload's identity to specific operations and resources. Broad policy is a finding only when lower-trust input can reach an unauthorized action.
|
||||
- Ingress, proxies, service mesh, metadata services, and admission policy are real boundaries, but only count a control when its configuration and attachment are visible.
|
||||
- Secret references are not secret disclosure. Require a lower-trust reader, output, artifact, log path, or unsafe fallback.
|
||||
- Use `confirmed` for active in-repo configurations and local rendering/policy validation. Use `needs_validation` for account policy, network attachment, runtime admission, hosted metadata, or drift that needs owner observation.
|
||||
```
|
||||
|
||||
## Workload identity and IAM attack classes (subagent_type: `general`)
|
||||
|
||||
**Workload identity overreach**
|
||||
A workload, pod, function, edge worker, or node identity can act on tenants, accounts, resources, or APIs beyond its role, and untrusted request or job input selects that target. Review cloud policy conditions, resource patterns, service-account attachment, namespace mapping, and fallback credentials.
|
||||
|
||||
**Cross-account or cross-tenant role confusion**
|
||||
Role assumption, external IDs, token exchange, workload federation, or resource policies accept identity claims not bound to the intended source account, audience, repository, namespace, or workload. Establish both trust policy and caller-controlled claim.
|
||||
|
||||
**Application authorization delegated to cloud metadata**
|
||||
An app trusts caller-supplied identity headers, tags, labels, account IDs, or resource metadata without verifying they came from the cloud control plane or a trusted proxy. Cloud IAM and application authorization are separate checks.
|
||||
|
||||
## Ingress, network, and control-plane attack classes (subagent_type: `general`)
|
||||
|
||||
**Unexpected service or management-plane reachability**
|
||||
An ingress, service, listener, security group, load-balancer annotation, port mapping, or server bind exposes an admin, debug, metrics, node, control-plane, or internal API to a lower-trust network. Missing network controls alone are `needs_validation`; a repository-controlled public route to a sensitive handler can be `confirmed`.
|
||||
|
||||
**Trusted-proxy and mesh identity bypass**
|
||||
A backend accepts forwarded identity, mTLS subject, or authorization metadata from peers outside the intended ingress/sidecar, or an alternate port and health/legacy path bypasses the mesh. Verify header stripping, peer reachability, and fail-open behavior when the proxy is absent.
|
||||
|
||||
**Metadata and internal-service reachability**
|
||||
An untrusted URL, destination, or protocol selection reaches instance/container metadata, control-plane sockets, or internal APIs with workload credentials. Trace URL parsing and redirect handling under `ATTACK-CLASSES.md`; here establish deployed network, metadata-version, and identity boundaries.
|
||||
|
||||
## Container and orchestration attack classes (subagent_type: `general`)
|
||||
|
||||
**Host or control-plane capability exposure**
|
||||
A lower-trust workload can select privileged mode, capabilities, host namespaces, host paths, device mounts, container runtime sockets, or service-account tokens that cross into node/control-plane authority. Bare absence of seccomp or read-only filesystem is hardening unless a reachable operation crosses that boundary.
|
||||
|
||||
**Admission and policy path inconsistency**
|
||||
One deployment route enforces image identity, namespace, resource, secret, or privilege policy while another controller, job, upgrade, restore, or compatibility path does not. Confirm the alternate route and resulting deployed object.
|
||||
|
||||
**Namespace and label trust confusion**
|
||||
Network, admission, secret, or workload-identity policy relies on labels, annotations, names, or namespaces that a less-trusted principal can set. Compare who controls selectors with what authority matching grants.
|
||||
|
||||
## Configuration and secret lifecycle attack classes (subagent_type: `general`)
|
||||
|
||||
**Security-control precedence drift**
|
||||
Development values, chart defaults, environment variables, command-line flags, feature gates, sidecar injection, or per-region overlays disable authentication, transport security, tenant scoping, or audit policy in a deployed environment. Render the final configuration for each maintained deployment, not just the base file.
|
||||
|
||||
**Secret exposure across workload boundaries**
|
||||
Secrets enter logs, crash reports, process arguments, shared environment, broad volumes, build outputs, service discovery, or read APIs accessible to another workload or tenant. Check secret type and authority; a public endpoint or key ID is not a credential.
|
||||
|
||||
**Credential renewal and outage fallback**
|
||||
Failure to mount, refresh, rotate, or revoke a workload credential causes stale credentials to remain active or an app to accept a less trusted identity mode. Review startup, readiness, reconnect, and cached-client behavior.
|
||||
|
||||
## Managed storage, events, and edge attack classes (subagent_type: `general`)
|
||||
|
||||
**Object and signed-URL policy confusion**
|
||||
Bucket/container policy, object keys, CDN origins, or signed URLs fail to bind principal, operation, object namespace, audience, or expiry. Review list/version operations and write paths as well as reads.
|
||||
|
||||
**Event-source identity confusion**
|
||||
A function or worker trusts event body fields as source identity without validating provider-signed envelope, subscription/topic, account, region, and replay state. Compare push, pull, retry, and dead-letter paths.
|
||||
|
||||
**Edge/runtime boundary mismatch**
|
||||
An edge or serverless runtime assumes a secret, API, filesystem, isolation, or tenant policy that differs from the origin runtime, and fallback to origin changes authority or cache behavior. Confirm which configuration selects each path.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Render every maintained environment and make a matrix of external port, workload identity, network peers, mounted secrets, and cloud resources. Differences require an owner or policy explanation.
|
||||
- Follow a lower-trust request, object, label, or event into cloud policy. Show which workload credential performs the final operation and what condition should scope it.
|
||||
- Diff normal deploy, migration, restore, node maintenance, failover, and local/emulator paths. Review behavior when mesh, admission, identity, secret, or policy service is unavailable.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Establish the active source path and effective deployment object; otherwise use `needs_validation` and state which rendered manifest or owner-observed attachment is missing.
|
||||
2. Name the lower-trust caller/workload, cloud or application identity, controllable selector, affected resource, and unauthorized operation or disclosure.
|
||||
3. Verify provider and orchestrator defaults at the pinned version. Do not assume a public IP, reachable metadata service, permissive firewall, or absent admission attachment.
|
||||
4. Local validation may render templates, evaluate policy, inspect container/user namespaces in an isolated fixture, or run an emulator with dummy identities. Do not probe live endpoints or alter shared cloud resources.
|
||||
5. Return `confirmed` only with a complete active source trace and concrete boundary result. Return `needs_validation` with the exact deployed policy, identity attachment, overlay, network, or drift observation needed.
|
||||
@@ -0,0 +1,84 @@
|
||||
# Data Isolation and Lifecycle Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when the target stores multi-tenant or access-controlled data, derives search/index/cache/analytics copies, issues object links, exports or restores records, migrates schemas, or promises deletion, revocation, and retention behavior. This domain follows one data item through every copy and state transition. Use `ATTACK-CLASSES.md` for endpoint-level access control and `CLOUD-AND-DEPLOYMENT.md` for provider-level storage policy.
|
||||
|
||||
Split large targets by primary storage, cache/search, object/blob storage, analytics/logging, export/backup, deletion/revocation, and migration.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- A tenant or owner field on a record is not isolation. Find the query, key, path, policy, or row-level control that enforces it for each read and write path.
|
||||
- Trace derived copies. Sanitized primary data can become unsafe in search, cache, analytics, export, previews, logs, replicas, and backups with different ACL and retention rules.
|
||||
- Deletion and revocation are lifecycle contracts. Check current, historical, cached, indexed, exported, restored, and queued copies within the product's stated boundary.
|
||||
- Privacy or retention preference is not automatically a security vulnerability. Require an explicit data-access boundary or deletion/revocation guarantee and an unauthorized reader or later operation.
|
||||
- Use `confirmed` for complete source-visible lineage and bounded dummy-tenant tests. Use `needs_validation` when external storage policy, retention, CDN behavior, replica lag, or backup access is unavailable.
|
||||
```
|
||||
|
||||
## Tenant and object-isolation attack classes (subagent_type: `general`)
|
||||
|
||||
**Missing tenant or owner enforcement**
|
||||
A read, update, delete, list, count, or bulk query identifies an object without binding it to the authenticated tenant/owner, or trusts body fields to supply that identity. Compare direct lookup, nested relationship, background, admin, import, and legacy paths.
|
||||
|
||||
**Composite-key and namespace collision**
|
||||
Cache keys, object paths, database uniqueness, search document IDs, temporary files, or deduplication keys omit tenant or environment. Two principals can overwrite or retrieve the same logical key even though application records carry separate owners.
|
||||
|
||||
**Policy and query disagreement**
|
||||
Row-level policy, ORM default scopes, authorization filters, and raw/bypass clients apply different predicates. Check joins, aggregates, aliases, views, transactions, `unscoped` or service clients, and error paths where context is missing.
|
||||
|
||||
**Blob and signed-reference overreach**
|
||||
Object keys, attachment IDs, version IDs, shared links, or signed URLs permit operations or namespaces beyond the issuing principal's access, or remain valid after the underlying ACL changes. Bind operation, exact object/version, audience, expiry, and tenant.
|
||||
|
||||
## Derived-data and disclosure attack classes (subagent_type: `general`)
|
||||
|
||||
**Search, cache, and index ACL drift**
|
||||
A primary record's ACL or lifecycle changes without invalidating a searchable, cached, embedded, thumbnail, RSS, preview, or index copy. Validate filtering at retrieval time as well as document ingestion and invalidation.
|
||||
|
||||
**Analytics, logs, traces, and diagnostics as alternate readers**
|
||||
Private content or credentials are emitted into systems with broader access, longer retention, or tenant mixing. Confirm the data class and realistic reader; field names, public identifiers, and operator-only content under intended policy are not enough.
|
||||
|
||||
**Enumeration and aggregate oracles**
|
||||
Counts, filters, ordering, errors, unique constraints, timings, notification behavior, or existence checks disclose protected object or account state. Require a concrete confidential predicate and observable distinction, not general response variance.
|
||||
|
||||
## Export, backup, restore, and migration attack classes (subagent_type: `general`)
|
||||
|
||||
**Export and backup scope expansion**
|
||||
An export, snapshot, portability package, report, or backup includes other tenants, inaccessible object fields, soft-deleted data, secret values, or history above the requester's access. Check per-item authorization after selection and authorization to download the final artifact.
|
||||
|
||||
**Import and restore authority expansion**
|
||||
Restore/import bypasses owner, schema, ACL, uniqueness, or validation rules, overwrites existing resources, or recreates records in a tenant the requester cannot write. Validate archive contents as untrusted and authorize the resulting operation rather than trusting prior provenance.
|
||||
|
||||
**Migration default and ownership confusion**
|
||||
Old records lack tenant/ACL/lifecycle fields, incompatible IDs collide, or partial rollout makes new and old readers apply different defaults. Review backfill, dual-read/write, compatibility, rollback, and resumed-migration paths.
|
||||
|
||||
**Backup and replication boundary drift**
|
||||
Encryption keys, storage accounts, cross-region replicas, restoration environments, or support snapshots have broader identity or tenant scope than primary data. Source can confirm only in-repo policy; hosted access and retention require `needs_validation`.
|
||||
|
||||
## Deletion, revocation, and lifecycle attack classes (subagent_type: `general`)
|
||||
|
||||
**Soft-delete and tombstone bypass**
|
||||
Direct lookup, search, relation traversal, object link, background processor, or restore ignores the lifecycle predicate and returns or acts on a deleted/revoked record. Check whether soft-deleted identifiers can be re-registered before all references are gone.
|
||||
|
||||
**Stale authorization and derived copy use**
|
||||
Membership removal, ACL update, consent withdrawal, secret revocation, or role downgrade does not invalidate sessions, caches, subscriptions, jobs, or materialized data that continue to authorize future operations.
|
||||
|
||||
**Retention and queued-work overrun**
|
||||
Deletion completes in primary storage while queued processors, retries, exports, analytics, or generated artifacts recreate or retain the data beyond the promised boundary. Find idempotent deletion and tombstone propagation.
|
||||
|
||||
**Restore reintroduces invalid state**
|
||||
Backup, undo, undelete, or replica recovery restores data, credentials, memberships, or permissions that current policy no longer allows. Re-authorize restored state and reapply lifecycle changes made after the snapshot.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Pick one protected record and draw primary write, query, cache, index, event, export, backup, deletion, and restore paths. Mark principal and tenant at every edge.
|
||||
- Compare two dummy tenants through the same local service methods, then repeat after ACL change, deletion, account switch, and restore. Do not use real user data.
|
||||
- Start at bypass clients, background jobs, migrations, global uniqueness, and cache keys. These paths commonly omit request-scoped identity that interactive endpoints carry.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Name attacker or lower-trust principal, protected data/state, affected owner/tenant, alternate copy or operation, and unauthorized disclosure or mutation.
|
||||
2. Cite both intended source-of-truth policy and the path that omits or disagrees with it. Confirm another layer does not enforce the same tenant/lifecycle condition.
|
||||
3. Use local dummy tenants and non-sensitive fixtures to prove cross-scope access or stale lifecycle behavior. Stop at the minimum observable record or operation.
|
||||
4. If external cache, object storage, replicas, analytics, backup, or retention policy is required, classify `needs_validation` and state the owner-observed check.
|
||||
5. Return `confirmed` only with complete lineage and concrete boundary impact. Return `needs_validation` with the exact unresolved storage, ACL, invalidation, retention, or restore fact.
|
||||
@@ -0,0 +1,89 @@
|
||||
# Desktop, Mobile, and Local IPC Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when the target is a desktop or mobile app, privileged helper, updater, local daemon, webview host, deep-link handler, browser native-messaging host, or local IPC client/server. Relevant untrusted actors may be a downloaded document, remote web content, another local app, another OS user, a sandboxed process, or a lower-privilege account. State that starting capability instead of treating all local users as equivalent.
|
||||
|
||||
Use `CLIENT-SIDE.md` for browser-side webview behavior, `MEMORY-SAFETY-AND-BINARY.md` for native memory and loader safety, and `SUPPLY-CHAIN-AND-RELEASE.md` for update authenticity.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- Establish the realistic local or remote-content attacker: another app, another OS user, a sandboxed child, an untrusted document, or a remote origin. Self-harm within the same account and authority is not a boundary violation.
|
||||
- Paths, process names, bundle/package IDs, and claimed sender fields are not peer authentication. Use OS peer credentials, code identity, capability handles, or protected channel state.
|
||||
- The native bridge or helper must authorize each operation and final resource after parsing. A trusted UI or broker does not make attacker-influenceable arguments trusted.
|
||||
- OS sandbox, signing, entitlements, permissions, keychain ACLs, exported-component policy, and prompt behavior are real controls when pinned and visible.
|
||||
- Use `confirmed` for source evidence plus bounded local/emulator tests. Use `needs_validation` when signing, manifest merge, OS version, device policy, installer ACL, or packaging is required but not observable.
|
||||
```
|
||||
|
||||
## Deep-link, callback, and navigation attack classes (subagent_type: `general`)
|
||||
|
||||
**Custom-scheme and deep-link ambiguity**
|
||||
Another app or page can invoke a route that mutates state, imports data, completes authentication, or selects an account without a current-session and one-time callback binding. Review URI normalization, duplicate query fields, scheme/host/path matching, exported activity/handler policy, and stale/replayed links.
|
||||
|
||||
**App and account handoff confusion**
|
||||
OAuth, SSO, magic-link, invite, device pairing, passwordless, or payment callbacks return to the wrong installed app, profile, tenant, or pending transaction. Bind state to the initiating app identity, current session, account, provider, operation, and expiry.
|
||||
|
||||
**File-open and intent authority confusion**
|
||||
An associated file, share intent, drag/drop item, pasteboard/clipboard record, notification action, or open-file event triggers a privileged operation without confirming content type, sender trust where applicable, current user intent, and final target.
|
||||
|
||||
## Webview and native-bridge attack classes (subagent_type: `general`)
|
||||
|
||||
**Navigation-origin to bridge confusion**
|
||||
Remote or attacker-controlled frames can reach a JavaScript/native bridge intended only for packaged content. Validate origin at call time and after every navigation, redirect, subframe creation, popup, and error/fallback page. URL-prefix checks and initial-load checks are insufficient.
|
||||
|
||||
**Over-broad native bridge capabilities**
|
||||
Web content can select arbitrary files, commands, IPC methods, credentials, or system actions through a generic bridge. Check method allowlists, normalized arguments, user/tenant authority, gesture/confirmation requirements, and return-value disclosure.
|
||||
|
||||
**Webview file and universal access**
|
||||
Remote content can read app-local files, privileged custom schemes, or internal origins because file access, universal access, mixed content, debug interfaces, or custom protocol handlers join origins unexpectedly. Missing a restrictive setting without reachable protected content is hardening.
|
||||
|
||||
## Local IPC and exported-component attack classes (subagent_type: `general`)
|
||||
|
||||
**IPC peer-authentication gaps**
|
||||
Unix sockets, named pipes, XPC, Binder, D-Bus, native messaging, RPC, shared memory, or loopback listeners accept a lower-trust peer without checking OS credentials, code identity, sandbox token, or channel ownership. Require a meaningful method or disclosure behind the channel.
|
||||
|
||||
**Claimed principal versus channel identity**
|
||||
The authenticated process/channel belongs to one app or user, but request fields select another user, tenant, profile, or capability. Bind each method and resource to the peer credential rather than a caller-declared identifier.
|
||||
|
||||
**Exported service, activity, receiver, or provider overreach**
|
||||
A mobile component or local automation endpoint is externally invokable and performs an operation intended for the app itself. Review final merged manifests, intent filters, permission/signature level, path grants, and alternate aliases. Manifest status unknown after packaging requires `needs_validation`.
|
||||
|
||||
**IPC lifecycle and correlation confusion**
|
||||
Predictable request IDs, reused handles, stale channels, inherited descriptors, world-writable socket paths, or restart behavior lets one peer answer, cancel, or reuse another peer's operation. Review creation permissions and cleanup of socket files, locks, ports, and shared mappings.
|
||||
|
||||
## Privileged-helper and local-file attack classes (subagent_type: `general`)
|
||||
|
||||
**Privileged helper as confused deputy**
|
||||
A low-privilege caller can select a privileged command, file, service, user, or system setting without per-operation authorization. Review sudo/polkit/UAC/XPC helper rules and ensure the helper independently validates normalized arguments.
|
||||
|
||||
**Install, update, and repair path trust**
|
||||
A privileged installer/helper reads manifests, scripts, packages, symlinks, working directories, or repair state writable by a lower-trust actor after authorization. Bind authorization to immutable content and safe destination paths.
|
||||
|
||||
**Local file ownership and TOCTOU**
|
||||
The app checks a file/path then follows replacement, symlink, mount, or case/normalization changes during a privileged read/write. Use descriptor-relative operations and verify final ownership. Focus `MEMORY-SAFETY-AND-BINARY.md` on parsing after the file is opened.
|
||||
|
||||
**Credential-store and local-secret boundary mismatch**
|
||||
A keychain/keystore item, token file, backup, log, clipboard, notification preview, or local config is readable by another app/profile/user with less authority. Plaintext readable only by the same intended OS account is not automatically a vulnerability; state the lower-trust reader and credential power.
|
||||
|
||||
## Application-state and device-lifecycle attack classes (subagent_type: `general`)
|
||||
|
||||
**Account switch, logout, and device restore leakage**
|
||||
Cached data, background tasks, widgets, notifications, local databases, webview storage, or biometric approvals survive logout/account change and appear under a later account. Review backup/restore and multi-profile behavior.
|
||||
|
||||
**Pending-action and user-presence confusion**
|
||||
Notification, widget, shortcut, share sheet, biometric prompt, or deferred operation authorizes a different action than displayed, executes after expiry, or uses another profile's pending state. Bind confirmation to normalized action, resource, account, and current foreground state.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Enumerate every process, app component, local endpoint, URI scheme, file association, webview origin, and helper. Record OS identity, runtime privilege, caller, and callable operation.
|
||||
- Read final packaging inputs: merged manifest, entitlements, installer rules, native-messaging registration, protocol handlers, and ACL creation. Source declarations can be overwritten downstream.
|
||||
- Validate with dummy profiles and non-sensitive local fixtures on an isolated machine/emulator. Do not interact with other users' apps, credentials, or production services.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Name the attacker starting capability, OS/app principal crossed, entry channel, accepted argument or state, and unauthorized operation or disclosure.
|
||||
2. Confirm OS sandbox, peer credential, signing, entitlement, permission, user-consent, and installer controls that apply. Unknown packaging/runtime facts require `needs_validation`.
|
||||
3. For webview bridges, cite both navigation/origin control and privileged native sink. For IPC, cite peer authentication and per-resource authorization. For helpers, verify final normalized destination.
|
||||
4. Keep local tests bounded and use dummy content/accounts. Stop after proving the boundary result; do not extend proof into persistence or broader system modification.
|
||||
5. Return `confirmed` only with a complete source and local evidence chain. Return `needs_validation` with the exact OS, manifest, signing, ACL, or device-lifecycle fact required.
|
||||
251
.agents/skills/security-audit/HUNTING.md
Normal file
251
.agents/skills/security-audit/HUNTING.md
Normal file
@@ -0,0 +1,251 @@
|
||||
# Vulnerability Hunting
|
||||
|
||||
### Phase 2: Run coverage-led hunting waves
|
||||
|
||||
The parent assigns `planned` ledger units to `general` agents. Use enough focused hunters to cover the units without combining unrelated boundaries. One hunter may own closely related units in one subsystem; no unit may be silently unassigned because of an agent-count limit — a unit the budget cannot reach is explicitly `deferred` with reason `budget_cannot_reserve_critics_and_validation`.
|
||||
|
||||
When a budget or profile caps hunter count, assign units in priority order and record the ordering rationale in the ledger. Rank by: (1) unauthenticated or lowest-trust entry surfaces before authenticated ones; (2) boundaries protecting the most valuable resources (credentials, cross-tenant data, code execution, release authority); (3) prior-run gaps, revalidation targets, and changed source before same-source re-passes; (4) units whose class historically yields confirmed findings for this target type over speculative ones. Ties break lexicographically by `coverage_id` so runs stay deterministic.
|
||||
|
||||
Before launch, the parent changes assigned units to `in_progress`, sets a canonical lowercase `agent_id`, and creates that agent's `scratch/` and parent-owned `artifacts/`. Hunters read source and parent-provided context, write only to their unique `scratch/`, and return one structured result through the Task tool. They never write retained artifacts or edit target source, `architecture.md`, `coverage-ledger.json`, `findings.json`, or another agent's files.
|
||||
|
||||
## Required hunter prompt
|
||||
|
||||
Every hunter prompt contains these parts in this order:
|
||||
|
||||
1. A two-sentence role preamble: the hunter's goal is to find source-grounded security invariant failures in its assigned units, and it must return exactly one JSON object matching the structured-result contract at the end of this prompt.
|
||||
2. `architecture.md` verbatim.
|
||||
3. Assigned coverage IDs, subsystem, boundary, repository-relative starting paths, and each unit's assignment block map from `coverage-ledger.json`.
|
||||
4. The exact selected blocks, copied verbatim: each selected ordinary attack-class block from `ATTACK-CLASSES.md`, and from each selected companion its `Core discipline`, each chosen attack-class subsection, `Universal moves`, and `Validation rules`. Ordinary blocks are self-contained and carry no companion-style `Core discipline`, `Universal moves`, or `Validation rules` sections. Do not send block or companion names alone.
|
||||
5. Explicit excluded ordinary and companion blocks with a reason for each exclusion.
|
||||
6. The core hunting method below, followed by the promotion procedure block.
|
||||
7. The core validation rules below.
|
||||
8. Carried same-source prior confirmed exclusions, each limited to fingerprint, title, and root cause, plus peer-owned current coverage IDs that this hunter must not duplicate.
|
||||
9. The unique scratch/artifact paths, safe agent ID, predeclared promotion allowlist and byte limits, and the structured-result contract, including the Structured hunter result block below and the `confirmed` and `needs_validation` branches of `report-schema.json` copied verbatim.
|
||||
|
||||
A prompt may select several companion blocks when the same path crosses several domains. Keep their constraints together. Scope is the hunter's coverage obligation, not permission to duplicate excluded work. If an unexpected different boundary appears, return it under `uncovered` so the parent creates a stable ledger unit and assigns it in the next wave.
|
||||
|
||||
#### Core hunting method — include in every hunter prompt
|
||||
|
||||
```text
|
||||
## Defensive vulnerability-finding method
|
||||
|
||||
Your goal is to find source-grounded security invariant failures and the smallest fix,
|
||||
not to expand harm beyond the boundary result. Stay within source review and bounded local execution.
|
||||
Do not contact deployed endpoints, provider APIs, registries, identity systems,
|
||||
message brokers, shared services, or other users. Use local dummy data only.
|
||||
|
||||
READ THE CODE AT DEPTH. Follow each assigned input through parsing, identity,
|
||||
authorization, normalization, state, derived copies, and the final sink. Read sibling,
|
||||
legacy, batch, retry, cancellation, migration, and error paths that produce the same
|
||||
effect. Compare sibling controls for equivalence, not only presence, and compare what
|
||||
one component guarantees with what the next component assumes.
|
||||
|
||||
WORK FROM A CONCRETE INVARIANT:
|
||||
1. Name the lower-trust principal and starting capability.
|
||||
2. Name the accepted value, action, state transition, or resource selector.
|
||||
3. Locate the control that should reject, bind, isolate, limit, or revoke it.
|
||||
4. Trace the exact source path after that decision.
|
||||
5. Stop at the smallest affected dummy record, wrong return value, process-integrity
|
||||
effect, or locally observable shared-resource effect.
|
||||
6. State a source-level change and regression case that enforce the invariant.
|
||||
|
||||
DEPTH BOUND: trace only paths that can reach your assigned boundary or whose
|
||||
guarantees that boundary relies on. Stop a line of investigation as soon as the
|
||||
invariant is settled either way, and record the result in your structured output —
|
||||
a covered, candidate, or blocked disposition, or an `uncovered` entry — instead of
|
||||
continuing to search.
|
||||
|
||||
TEST SAD PATHS AND DISAGREEMENTS. Check absent, empty, zero, negative, maximum,
|
||||
over-limit, duplicate, mixed encoding, stale, revoked, reordered, concurrent,
|
||||
partially migrated, failed dependency, and rollback state only where the interface
|
||||
accepts them. Compare canonicalization and units at every parser or policy handoff.
|
||||
For multi-step issues, treat each output as a prerequisite and do not assume a later
|
||||
boundary. If any prerequisite is not established, record a blocker.
|
||||
|
||||
When a proposed high or critical candidate reveals a reusable root cause, search paths
|
||||
owned by the assigned coverage IDs for lexical, structural, and logical variants.
|
||||
Consolidate the same root cause, but establish each variant's conditions and impact
|
||||
independently. Do not investigate peer-owned units. Return a variant with no current
|
||||
coverage unit as `uncovered`.
|
||||
|
||||
USE THE NARROWEST LOCAL CHECK THAT SETTLES THE CLAIM. Target-controlled builds,
|
||||
tests, processes, browsers, emulators, fuzzers, and fixture processing may run only
|
||||
inside the parent-approved OS-enforced sandbox. It must disable external networking,
|
||||
start from an empty allowlisted environment, expose target and tools read-only, permit
|
||||
writes only to your scratch directory, and apply low CPU, memory, process, file-size,
|
||||
disk, and wall-clock limits. Isolated loopback is allowed only for a local fixture.
|
||||
If any control is unavailable, do not execute: return needs_validation with that exact
|
||||
blocker. Prefer an existing unit test, minimal function harness, dummy-tenant service
|
||||
call, small malformed fixture, deterministic race schedule, or locally rendered policy.
|
||||
Do not install or fetch tools.
|
||||
|
||||
Record the exact input, command, limits, and minimum result. For the environment,
|
||||
record only allowlisted variable names and safe non-secret values needed to reproduce
|
||||
the check. Never capture the ambient environment, inherited variables, credentials,
|
||||
authentication state, or unrelated host paths. The target-controlled process writes
|
||||
only in scratch. After the sandbox and all its processes terminate, only trusted
|
||||
parent-side code may promote predeclared scratch-relative files, following the
|
||||
promotion procedure block included verbatim in this prompt. You and target code never
|
||||
write retained artifacts. If promotion is unavailable or fails for decisive evidence,
|
||||
return needs_validation with the exact promotion blocker.
|
||||
Never stress availability, invoke a live target, use a real credential, publish an
|
||||
artifact, or continue past the minimum observed effect.
|
||||
|
||||
A deployment, browser, provider, broker, OS, proxy, package, secret, or identity fact
|
||||
outside source is not proof either way. If one such fact is decisive, return a
|
||||
needs_validation record with the exact missing observation and safe owner-observed check.
|
||||
```
|
||||
|
||||
#### Promotion procedure — copy this promotion procedure verbatim into every hunter prompt
|
||||
|
||||
```text
|
||||
Artifact promotion procedure (trusted parent-side code only):
|
||||
Reference only for you: the parent performs these steps; you never perform them.
|
||||
|
||||
Before execution, the parent opens and retains trusted, non-inheritable directory
|
||||
descriptors for the agent's scratch/ and artifacts/ roots, and records an allowlist
|
||||
of expected scratch-relative artifact files plus explicit per-file and cumulative
|
||||
byte limits. Never pass those descriptors to the agent or sandbox. After the sandbox
|
||||
and all its processes terminate, trusted parent-side code promotes each allowlisted
|
||||
file separately:
|
||||
|
||||
1. Validate the declared relative path: reject absolute, empty, `.`, `..`, or
|
||||
symlinked components.
|
||||
2. Walk each parent component from the retained scratch-root descriptor with
|
||||
no-follow directory-relative operations; never reopen by path.
|
||||
3. Open the leaf no-follow and nonblocking.
|
||||
4. Verify with `fstat` that it is a regular file with link count exactly one and
|
||||
within the recorded per-file and cumulative byte limits.
|
||||
5. Enforce those limits again while reading from that descriptor.
|
||||
6. Copy exactly the verified size, repeat `fstat`, and reject a changed identity,
|
||||
type, link count, or size.
|
||||
7. For the destination, walk every parent component from the retained
|
||||
artifacts-root descriptor with no-follow directory-relative operations; require
|
||||
each existing component to be a real directory, and create any missing directory
|
||||
exclusively before reopening and verifying it no-follow.
|
||||
8. Create the leaf exclusively without following links, verify that the opened
|
||||
destination is a regular file with link count exactly one, and copy from the
|
||||
verified source descriptor without reopening either path.
|
||||
9. Use equivalent race-safe APIs on non-POSIX systems.
|
||||
10. Never recursively copy or glob scratch, extract an archive into artifacts, or
|
||||
open or promote a symlink, FIFO, socket, device, directory, hard-linked file,
|
||||
changing file, or file that exceeds its bound.
|
||||
11. If any check is unavailable, cannot be enforced, or fails, discard the scratch
|
||||
entry; if it is decisive evidence, retain `needs_validation` with the exact
|
||||
promotion blocker.
|
||||
```
|
||||
|
||||
#### Core validation rules — include in every hunter prompt
|
||||
|
||||
```text
|
||||
## Candidate gate
|
||||
|
||||
1. A candidate needs a complete repository-relative source trace and evidence for the
|
||||
claimed root cause, including the strongest source-visible control.
|
||||
2. A proposed confirmed record needs a bounded local observed result, meaningful impact
|
||||
across a stated boundary, complete conditions, and no visible preventing layer.
|
||||
3. Do not strengthen a crash into code execution, ordinary work into shared availability,
|
||||
or a same-principal action into privilege gain.
|
||||
4. If a required fact is not source-visible or locally observable, use
|
||||
needs_validation. Name exact blockers; do not give it severity or speculative completion.
|
||||
5. A missing best practice with no affected principal/resource is excluded or hardening,
|
||||
not a finding. A candidate disproved by source is not needs_validation.
|
||||
6. Use the same source-derived fingerprint for the same root cause in every state.
|
||||
It must match `^[A-Za-z0-9][A-Za-z0-9._:/@+-]*$` and must not include a line,
|
||||
wave, agent, severity, or verdict.
|
||||
7. Return an empty candidate array when nothing survives these gates.
|
||||
```
|
||||
|
||||
## Local validation boundaries
|
||||
|
||||
Local execution is for confirmation, not impact expansion:
|
||||
|
||||
- **Allowed only in the required OS sandbox:** offline builds with present dependencies; isolated-loopback processes using dummy state; unit and integration tests; small fixture processing; sanitizers; bounded fuzz/regression tests; deterministic concurrency checks; local browser/emulator tests with dummy accounts; rendered manifests and policy evaluation with dummy identities; mocked external or paid calls.
|
||||
- **Disallowed:** live or deployed traffic; requests to services not started for this isolated check; network dependency installation; real accounts or credentials; production data; shared queues, cloud resources, runners, registries, signing or release services; publishing; stress, saturation, or cost generation; any work after the minimum dummy-data boundary result.
|
||||
|
||||
The sandbox starts with an empty environment, gives target code no external network or host writable path, and enforces explicit low resource and time limits for every check, not only checks expected to be expensive. Scratch output remains target-controlled after exit. Promote it only with the no-follow, path-confined, regular-file, bounded-size host procedure in `SKILL.md`. Missing any sandbox or promotion capability does not erase a source-grounded candidate; represent the exact blocker in `needs_validation`.
|
||||
|
||||
## Structured hunter result
|
||||
|
||||
Return exactly one JSON object, with no surrounding prose:
|
||||
|
||||
```json
|
||||
{
|
||||
"units": [
|
||||
{
|
||||
"coverage_id": "one assigned ID",
|
||||
"disposition": "covered|candidate|blocked",
|
||||
"reviewed_paths": ["repo/relative/path"],
|
||||
"checks": [
|
||||
{
|
||||
"agent_id": "canonical owner of this check",
|
||||
"reviewed_paths": ["repo/relative/path owned by this check"],
|
||||
"invariant": "specific control checked for this unit",
|
||||
"method": "source|local",
|
||||
"result": "what source or the bounded check established",
|
||||
"artifact": "agents/<agent-id>/artifacts/file for local, null for source"
|
||||
}
|
||||
],
|
||||
"candidate_fingerprints": [],
|
||||
"unresolved": []
|
||||
}
|
||||
],
|
||||
"candidates": [],
|
||||
"hardening": ["concrete non-finding note"],
|
||||
"uncovered": [
|
||||
{
|
||||
"surface": "...",
|
||||
"boundary": "...",
|
||||
"subsystem": "...",
|
||||
"attack_class": "...",
|
||||
"starting_paths": ["repo/relative/path"],
|
||||
"reason": "why this needs its own deterministic coverage unit"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Each `candidates` entry is schema-shaped except that it uses `proposed_verdict` in place of `verdict`:
|
||||
|
||||
- `proposed_verdict: "confirmed"`: include every field required by the `confirmed` branch of `report-schema.json` other than `verdict`: `fingerprint`, title, description, `root_cause`, `intended_behavior`, ordered `trace`, `evidence`, `conditions`, target-neutral `execution`, `remediation`, `severity`, and `confidence`. The execution instructions describe only the bounded local check already performed. `payloads` holds the minimum test input, fixture, or native invocation. `observed_result` records actual local output. Overall severity must not exceed observed impact.
|
||||
- `proposed_verdict: "needs_validation"`: include every field required by that schema branch other than `verdict`: `fingerprint`, title, description, `claimed_root_cause`, ordered `trace`, `evidence`, nonempty `blockers`, and `validation_plan` with at least one applicable `local` or `deployment` step. Do not invent an inapplicable context. Do not include severity, execution, remediation, reason, or a confirmed `root_cause`. `deployment` is an owner-observed check, not a request to probe a live target.
|
||||
|
||||
Every assigned coverage ID appears exactly once in `units`. A `covered` unit needs an owner, nonempty `reviewed_paths` and `checks`, no unresolved fact, and no candidate. A `candidate` unit has the same owned evidence and is the only state that carries linked fingerprints. A `blocked` unit is an owned partial review with nonempty paths, checks, and unresolved facts but no fingerprint. All source paths are repository-relative, never absolute or traversal paths. A trace with several entries begins at `entrypoint`, ends at `sink`, and labels intermediate steps `propagation`. Every check has its own canonical lowercase `agent_id` and nonempty `reviewed_paths`; the unit-level list is exactly the union of those owned paths. A `source` check uses `artifact: null`. A `local` check uses one successfully parent-promoted regular file beneath `agents/<check.agent_id>/artifacts/`; this permits a verifier to add independently owned evidence without taking ownership from the hunter. Never link scratch, an output-root file, or another check owner's artifact.
|
||||
|
||||
## Parent consolidation and ledger update
|
||||
|
||||
The parent validates each unit result, maps it to exactly one assigned `coverage_id`, and updates only that ledger unit. Reject duplicate or absent IDs, unsafe unit or check agent IDs, source checks with artifacts, and local artifacts that trusted parent-side code did not promote into the check owner's artifacts subtree. Copy the unit's `reviewed_paths`, its `checks` into the unit's `local_checks`, linked artifact paths, candidate fingerprints, and unresolved facts into the ledger. Retain each hunter's `hardening` list in a parent bookkeeping field on the relevant units (outside the semantic fields) so Phase 6 can report it. A failed, malformed, or unsupported conclusion leaves that unit `planned` for reassignment. Untouched budget/profile units become unassigned `deferred` units with empty evidence and a reason; do not hide partial evidence in `deferred`. Run `validate-coverage-ledger.cjs` after the update; an invalid ledger cannot drive another assignment. This per-unit contract allows one hunter to close one unit while returning a candidate or blocker for another.
|
||||
|
||||
Consolidate candidate entries by fingerprint and then by root cause. One root cause that exposes several entry paths or effects is one candidate with the strongest complete trace. Related but independent missing controls use separate fingerprints. Record duplicate fingerprints in the relevant ledger unit and do not send duplicate candidates to validation.
|
||||
|
||||
## Coverage-critic waves
|
||||
|
||||
Immediately after each hunter wave, spend the reserved invocation on one fresh `research` post-wave coverage critic. It receives `architecture.md`, the full coverage ledger including each assignment block map, current candidate fingerprints and states, and the prior-ledger gap summary. It reads source but does not write or run targets. Require exactly this JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"missing_units": [
|
||||
{
|
||||
"surface": "...",
|
||||
"boundary": "...",
|
||||
"subsystem": "...",
|
||||
"attack_class": "...",
|
||||
"starting_paths": ["repo/relative/path"],
|
||||
"selected_companion_blocks": ["FILE.md#section"],
|
||||
"excluded_blocks": [{"block": "FILE.md#section", "reason": "..."}],
|
||||
"reason": "source-backed coverage gap"
|
||||
}
|
||||
],
|
||||
"reassign_ids": ["existing-id-that-did-not-close"],
|
||||
"resolved_prior_leads": ["fingerprint"],
|
||||
"stop": false
|
||||
}
|
||||
```
|
||||
|
||||
The critic checks for unmapped entry points, unchecked parallel paths, missing lifecycle modes, selected companion classes without a unit, unjustified exclusions, units closed without paths/checks, and prior `needs_validation` or changed-source gaps that no unit addresses. It proposes coverage, not findings. `stop` is the critic's own assessment: `true` only when it accepts no `missing_units` and no `reassign_ids`; the parent's loop condition below, not `stop` alone, decides whether another wave runs. For each fingerprint in `resolved_prior_leads`, the parent marks the linked unit or prior-lead entry resolved and records the critic's source-backed reason.
|
||||
|
||||
The parent rejects proposed units outside the review scope or source/local boundary, derives canonical IDs for accepted units, and deduplicates them against current units. A prior same-source completed unit may supply evidence; prior `deferred`, `blocked`, `out_of_scope`, or changed-source units become current work and never suppress an accepted unit. Fail rather than merge a canonical ID collision. For each legitimate `reassign_id` with live `blocked`, `covered`, or `candidate` evidence, append that exact terminal record to the unit's `attempts` with the critic's source-backed `reassignment_reason`. Preserve its owner, checks, artifacts, fingerprints, and unresolved facts in that archive. Increment the live `wave`; the next hunter must be a fresh owner and receives an `in_progress` unit with empty live evidence. The hunter's terminal result writes only its new evidence into the live fields. Never copy an archived owner's checks or artifacts into the new live attempt. Sort IDs and validate the ledger before another assignment. In `standard` and `deep`, when the post-wave critic reports no accepted `missing_units` or legitimate `reassign_ids` and no `planned` units remain, spend the separately reserved invocation on a distinct final-clean critic. Complete coverage only when that critic also returns no accepted work. If it finds work, queue it and repeat the wave, post-wave critic, and final-clean process. If time or resources force an early stop, mark every untouched unit `deferred`, preserve the critic's reason, and disclose the gap in the report. Never use a silent wave or agent cap as evidence of complete coverage.
|
||||
|
||||
The run profile bounds this loop. A `quick` run has exactly one hunter wave followed by exactly one final critic pass. Add each accepted `missing_unit` to the current ledger and mark it `deferred` with reason `quick_profile_final_critic`. For each legitimate evidence-bearing `reassign_id`, archive the live terminal state in `attempts`, increment `wave`, and set the live state to unassigned `deferred` with empty evidence and reason `quick_profile_final_critic`. Do not launch a second hunter wave or another critic. In a scoped run, the critic still reports out-of-scope gaps it notices, but the parent records them as `out_of_scope` with the critic's reason instead of assigning them. The early-stop rule above is the same mechanism: `quick` is a pre-declared early stop, not evidence of complete coverage.
|
||||
|
||||
A budget bounds it the same way. Before assigning each wave, compare remaining budget against its hunter count, the validation reserve, the immediate post-wave critic, and the retained final-clean critic (`quick` reserves only its single final post-wave critic). Shrink the hunter wave to fit, taking units in priority order. If those mandatory reserves do not fit, launch no hunter from that wave and mark its planned units `deferred` with reason `budget_cannot_reserve_critics_and_validation`. Critic-proposed units enter the same ranked queue rather than extending the budget. If surviving candidates exceed the validation reserve, follow the incomplete-run rule in `SKILL.md`: stop hunting, validate in fingerprint order, retain unvalidated units as unresolved candidates, and never present them as findings.
|
||||
101
.agents/skills/security-audit/MEMORY-SAFETY-AND-BINARY.md
Normal file
101
.agents/skills/security-audit/MEMORY-SAFETY-AND-BINARY.md
Normal file
@@ -0,0 +1,101 @@
|
||||
# Memory Safety, Binary, and Kernel Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when the target processes untrusted bytes in a memory-unsafe or privileged context: C/C++/Objective-C, Rust `unsafe`, FFI, kernel modules and drivers, parsers and decoders, network daemons, firmware, binary loaders, language runtimes, and JITs. Use `PROTOCOLS-RPC-AND-MESSAGING.md` for protocol authorization and state-machine logic, and this file for process integrity, memory safety, ABI boundaries, and loader behavior.
|
||||
|
||||
Pick relevant classes from Phase 1 and split large targets by parser, allocator/lifetime, FFI, concurrency, loader, runtime, or privileged interface.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- Re-derive every bound and lifetime from attacker-controlled inputs and all callers. Validate against the worst accepted case, not a typical test vector.
|
||||
- A panic, sanitizer finding, or crash proves a defect only when a realistic untrusted input reaches it. Do not infer memory corruption, code execution, or shared availability impact from a label alone.
|
||||
- Validate in a local harness with sanitizers, deterministic concurrency tests, existing fuzz targets, and debugger-assisted fault classification. Stop after proving the violated invariant and observable impact; do not develop post-corruption techniques.
|
||||
- Assembly, JIT code, custom allocators, intra-object accesses, and foreign libraries can escape sanitizer coverage. Identify which relevant instructions are instrumented.
|
||||
- Classify as `confirmed` only after source evidence and bounded local validation establish the defect and effect. Use `needs_validation` when ABI, allocator, architecture, feature, deployment, or reachability facts remain unknown.
|
||||
```
|
||||
|
||||
## Bounds, integer, and representation attack classes (subagent_type: `general`)
|
||||
|
||||
**Out-of-bounds read or write**
|
||||
A length, offset, index, or terminator reaches a fixed or allocated buffer without a correct bound. Recalculate available headroom after prefixes, alignment, padding, and terminators. Check both source and destination capacity, and whether a short input is read before its declared length is trusted.
|
||||
|
||||
**Integer overflow, underflow, truncation, and signedness**
|
||||
Review attacker-controlled arithmetic before allocation, copy, loop, indexing, and pointer operations. High-hit patterns include `a - b` with `b > a`, `count * element_size`, additions near the type maximum, negative values converted to unsigned, 64-bit lengths narrowed to 32-bit fields, and sentinel values such as `-1` becoming a large size. Confirm which checked representation is later used.
|
||||
|
||||
**Unit and pointer-depth confusion**
|
||||
Code mixes bytes, elements, code units, pages, words, wire units, or nested pointer element sizes. Compare the unit at parse, validation, allocation, API boundary, and copy. A bounds check using the same wrong unit as the allocation is still wrong.
|
||||
|
||||
**Uninitialized or partially initialized data**
|
||||
A buffer, padding, struct field, or vector capacity is returned, compared, hashed, serialized, or passed across a trust boundary before initialization. Require an observable consumer and realistic output length; stack allocation by itself is not disclosure.
|
||||
|
||||
## Lifetime, type, and concurrency attack classes (subagent_type: `general`)
|
||||
|
||||
**Use-after-free, stale view, and double free**
|
||||
Owners are released while callbacks, wait queues, timers, iterators, borrowed slices, or cached raw pointers can still use them. Review every error, cancellation, close, and realloc path. For embedded notification anchors, each free path must drain or detach all observers.
|
||||
|
||||
**Type confusion and invalid downcast**
|
||||
A tag, vtable, union discriminator, object kind, or foreign handle is checked differently from the representation later read. Look for unchecked dynamic casts, stale tags after reuse, and serialized types whose validated element differs from the element consumed. Confirm a wrong-type read or write locally without extending the test beyond the violated invariant.
|
||||
|
||||
**Reference-count and ownership races**
|
||||
Non-atomic retain/release, a check followed by an unlocked use, or inconsistent ownership across threads can free or mutate an object during access. Compare fast, error, shutdown, and compatibility paths for the same lock and ownership rules.
|
||||
|
||||
**Shared-state races and TOCTOU**
|
||||
Concurrent parser streams, global caches, lazy initialization, signal handlers, and resource teardown can invalidate bounds, policy, or pointers established earlier. Verify the race with a repeatable local schedule, barrier, or thread sanitizer; a hypothetical interleaving without a security-relevant state transition remains `needs_validation`.
|
||||
|
||||
**Lock-order, deadlock, and starvation**
|
||||
Externally reachable operations acquire locks in inconsistent order or hold them across callbacks and blocking I/O. Report under availability only when bounded input can stop shared progress; otherwise record it for fixing as a concurrency defect.
|
||||
|
||||
## FFI and ABI attack classes (subagent_type: `general`)
|
||||
|
||||
**Pointer-length and ownership contract mismatch**
|
||||
Caller and callee disagree on who allocates, frees, pins, or mutates a buffer, how long a pointer remains valid, or whether a length is bytes or elements. Trace both sides of every `extern`, CGo/JNI/Python/native binding, and generated wrapper. Check null, zero length, aliasing, and callback retention.
|
||||
|
||||
**Layout, alignment, and enum disagreement**
|
||||
Foreign code receives a struct, bitfield, packed record, callback signature, integer width, enum, or calling convention that differs by architecture or build flag. Verify `repr`, packing, alignment, endianness, and ABI-specific types. An in-repo declaration mismatch can be confirmed locally; an opaque foreign implementation requires `needs_validation`.
|
||||
|
||||
**Unwind, exception, and thread-affinity violations**
|
||||
Exceptions or panics cross an ABI that forbids unwinding, callbacks run after teardown, or APIs requiring one runtime thread are invoked elsewhere. Review error conversion and cancellation. Confirm whether the process aborts or state is corrupted before assigning impact.
|
||||
|
||||
## Binary loading and runtime attack classes (subagent_type: `general`)
|
||||
|
||||
**Library, plugin, and executable search-order trust**
|
||||
A privileged process loads a library, plugin, runtime image, or helper from a path writable by a less-trusted principal, or resolves a bare name through an attacker-influenceable working directory or environment. Compare intended installation ownership with each fallback and compatibility search path. A user loading their own plugin into their own process is not a boundary violation.
|
||||
|
||||
**Missing artifact identity or signature binding**
|
||||
A loader verifies one file or metadata record but maps a different image because path resolution, file replacement, architecture slices, or embedded resources are not bound to the check. Supply-channel authenticity belongs in `SUPPLY-CHAIN-AND-RELEASE.md`; this class covers the local verification-to-map gap.
|
||||
|
||||
**Malformed binary metadata and relocation handling**
|
||||
Offsets, counts, sections, relocations, symbols, bytecode, or debug metadata are trusted before range, overlap, and representation checks. Test parsers with bounded local fixtures and sanitizers. Separate memory corruption from a safely rejected malformed file.
|
||||
|
||||
**JIT and generated-code consistency**
|
||||
Validator, interpreter, optimizer, and generated code disagree about types, bounds, side effects, or lifetime. Diff optimized and unoptimized paths using the same local input. Confirm a process-integrity effect; output variance that stays within language semantics is not a finding.
|
||||
|
||||
**Unload, reload, and teardown safety**
|
||||
Live function pointers, callbacks, worker threads, or data views survive module unload or runtime reset. Review shutdown and failed-load cleanup as closely as startup.
|
||||
|
||||
## Kernel and privileged-interface attack classes (subagent_type: `general`)
|
||||
|
||||
**User-copy bounds and repeated reads**
|
||||
A syscall, ioctl, driver, or kernel parser derives a trusted fact from user memory then reads the same mutable address again. Copy the full request once or revalidate the later copy. Also audit size, direction, and access checks at each user-copy primitive.
|
||||
|
||||
**Privileged object lifecycle and dispatch consistency**
|
||||
Externally reachable objects have unbalanced retain/release, teardown without observer drain, unchecked selector/table indices, or duplicated compatibility paths that omit a guard. Diff each dispatch and free path side by side.
|
||||
|
||||
**Under-authorized powerful interfaces**
|
||||
A device node, admin socket, helper, or management API validates shape but not the caller's authority over the resource. Establish actual interface ownership and reachability; permissions or sandbox policy outside the repository make this `needs_validation`.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Audit fixes and duplicated paths for the same source-to-sink shape. A check in one caller, architecture, protocol role, feature flag, or compatibility path does not protect its siblings.
|
||||
- Build a table for every parser or FFI boundary: accepted length/type, checked representation, allocation owner, consumer, thread, and teardown. Most native findings are one disagreement in that table.
|
||||
- Use existing corpora and small locally generated boundary fixtures. Save exact sanitizer/runtime output and the input property that triggers it; avoid large resource consumption and any live target.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Establish a realistic untrusted entry and exact operation that violates a bounds, type, lifetime, ABI, concurrency, loader, or authority invariant.
|
||||
2. Classify the observable effect: invalid read, invalid write, stale alias, wrong object, uninitialized output, unauthorized image load, deadlock, or safe process termination. Do not claim a stronger effect than observed.
|
||||
3. Run the narrowest local harness, existing test, sanitizer, or fuzzer needed to reproduce the effect. Verify sanitizer coverage of the faulting operation and record architecture/build conditions.
|
||||
4. For concurrency, use a deterministic schedule or sanitizer trace. For binary loading, prove the checked identity differs from the mapped identity and name the lower-trust writer.
|
||||
5. Return `confirmed` findings only with exact input, source trace, and observed result. Return `needs_validation` for a specific unresolved reachability, ABI, build, deployment, or runtime fact and state the bounded check needed.
|
||||
81
.agents/skills/security-audit/PROTOCOLS-RPC-AND-MESSAGING.md
Normal file
81
.agents/skills/security-audit/PROTOCOLS-RPC-AND-MESSAGING.md
Normal file
@@ -0,0 +1,81 @@
|
||||
# Protocols, RPC, and Messaging Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when the target uses gRPC, GraphQL transports, Cap'n Proto, Thrift, Protobuf, custom binary protocols, streaming RPC, webhooks, brokers, queues, pub/sub, or event buses. It covers peer identity, logical message interpretation, routing, replay, ordering, and delivery semantics. Use `MEMORY-SAFETY-AND-BINARY.md` for parser memory safety, `WEB-PROTOCOL-AND-AUTH.md` for HTTP framing, and `RESOURCE-EXHAUSTION-AND-AVAILABILITY.md` for availability impact.
|
||||
|
||||
Split large systems by producer/consumer pair, external/internal peer role, synchronous RPC, streaming, and asynchronous message path.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- "Internal" is not authentication. Name the peer identity at every hop and show how it becomes the application principal used for authorization.
|
||||
- Schema validation proves message shape, not provenance, resource authority, ordering, or safe values. Follow decoded fields to policy and side effects.
|
||||
- Broker guarantees and application guarantees differ. Write down retry, ordering, acknowledgement, deduplication, and transaction behavior before evaluating state changes.
|
||||
- Parser disagreement requires two concrete consumers, schema versions, or wire representations and one security-relevant divergent value.
|
||||
- Use `confirmed` for source-complete paths plus bounded local producer/consumer tests. Use `needs_validation` for broker ACL, service-mesh identity, topic attachment, or compatibility behavior outside the repository.
|
||||
```
|
||||
|
||||
## Framing, schema, and interpretation attack classes (subagent_type: `general`)
|
||||
|
||||
**Message boundary and canonicalization disagreement**
|
||||
Components disagree on length, compression, duplicate fields, unknown fields, encoding, numeric width, normalization, or envelope/body precedence. Compare generated and custom parsers, gateways, language bindings, and version converters. Confirm which principal, resource, or operation differs after decoding.
|
||||
|
||||
**Union, enum, and default confusion**
|
||||
Unknown variants, missing discriminators, zero values, default privileges, or compatibility mappings reach code that assumes a validated case. Review exhaustive dispatch, default branches, and how old consumers interpret newly added fields.
|
||||
|
||||
**Envelope and payload identity mismatch**
|
||||
Authorization uses trusted-looking routing or envelope metadata while the handler acts on a conflicting tenant, account, subject, object, or sender in the body. Identify which source is authoritative and ensure clients cannot override it.
|
||||
|
||||
## RPC identity and authorization attack classes (subagent_type: `general`)
|
||||
|
||||
**Interceptor and method-path inconsistency**
|
||||
An authn/authz interceptor applies to unary methods but not streams, reflection, health, gateway-transcoded paths, compatibility services, or individual stream messages. Compare every registration and route to the same operation.
|
||||
|
||||
**Peer identity to application-principal confusion**
|
||||
mTLS, workload identity, bearer metadata, forwarded identity, or broker credentials authenticate a channel, but a caller-controlled field selects the user or tenant. The channel identity and claimed principal must be bound by deterministic policy.
|
||||
|
||||
**Per-item and streaming authorization gaps**
|
||||
A stream, subscription, batch, or bulk message is authorized once, then later items name different resources or continue after role, membership, or token revocation. Re-check where scope can change and bind subscriptions to their original principal.
|
||||
|
||||
**Callback and reply-correlation confusion**
|
||||
Predictable, reused, or cross-tenant correlation IDs let a response, webhook, cancellation, or acknowledgment satisfy another caller's pending operation. Bind each outstanding request to authenticated peer, tenant, operation, and lifecycle.
|
||||
|
||||
## Broker and queue isolation attack classes (subagent_type: `general`)
|
||||
|
||||
**Topic, routing-key, and subscription scope gaps**
|
||||
A publisher or subscriber can select another tenant's topic, wildcard, consumer group, partition, reply queue, or dead-letter route. Check broker-enforced ACLs where visible and application-side namespace construction. Tenant text inside a payload is not isolation.
|
||||
|
||||
**Dead-letter, retry, and diagnostic disclosure**
|
||||
Messages routed to dead-letter queues, error topics, tracing, or operator views contain secrets or cross-tenant payloads accessible to a lower-trust consumer. Review policy and redaction at the failure path, not just normal delivery.
|
||||
|
||||
**Untrusted producer treated as control plane**
|
||||
A message body can declare itself an admin event, provider callback, replication record, or migration instruction without an independently authenticated producer and event type. Verify signatures and source/account/audience binding before privileged handling.
|
||||
|
||||
## Replay, ordering, and transaction attack classes (subagent_type: `general`)
|
||||
|
||||
**Duplicate delivery and idempotency gaps**
|
||||
Retries or redelivery repeat a side effect because deduplication is absent, occurs after mutation, or uses a key that collides across tenants or operations. Confirm the broker's delivery model and the side effect that is not naturally idempotent.
|
||||
|
||||
**Out-of-order and stale message acceptance**
|
||||
Older state, revoked membership, canceled work, or pre-step-up authorization arrives after newer state and overwrites it. Review sequence/version checks, tombstones, partition changes, and restore/replay workflows.
|
||||
|
||||
**Acknowledgment/commit ordering defects**
|
||||
Acknowledgment occurs before durable commit and loses security-relevant work, or commit happens before an unreliable acknowledgment and duplicates a mutation. Evaluate transactional outbox/inbox behavior and failure recovery.
|
||||
|
||||
**Partial multi-consumer transitions**
|
||||
Several consumers jointly implement one authorization or business transition, but retries and partial failure leave only a subset committed. Identify invariants that must become durable atomically or compensate with current authorization.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Draw producer → broker/transport → gateway → consumer → storage for each message family. At each hop record authenticated peer, authoritative tenant/resource fields, validation, and side effect.
|
||||
- Feed the same small fixture to every in-repo schema version or language binding. Test duplicate, missing, unknown, boundary, replayed, and reordered messages without producing load.
|
||||
- Compare normal, retry, dead-letter, replay, migration, reflection, stream, and gateway-transcoded routes. Security policy must survive transport changes.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Name the realistic producer or peer, accepted message, authenticated channel identity, affected principal/resource, and unauthorized mutation or disclosure.
|
||||
2. For disagreement claims, cite both parsers/consumers and the divergent decoded value. Safe rejection by either side prevents confirmation.
|
||||
3. For replay/order claims, establish actual delivery guarantees and reproduce the invariant failure with a bounded local/in-memory transport.
|
||||
4. For authorization and isolation, verify all interceptor, broker ACL, gateway, and consumer layers visible in source. External attachments make the candidate `needs_validation`.
|
||||
5. Return `confirmed` only with the complete message lifecycle and observed meaningful result. Return `needs_validation` with the exact broker, service identity, route, or delivery fact required.
|
||||
156
.agents/skills/security-audit/RECONNAISSANCE.md
Normal file
156
.agents/skills/security-audit/RECONNAISSANCE.md
Normal file
@@ -0,0 +1,156 @@
|
||||
# Reconnaissance
|
||||
|
||||
### Phase 1: Map the source and plan coverage
|
||||
|
||||
The parent initializes `run-metadata.json`, applies the strict pre-reconnaissance budget gate in `SKILL.md`, then creates agent scratch roots and the shared ledger before hunting. If the gate fails, record the incomplete status in metadata and launch no reconnaissance agent. Reconnaissance reads the target and locally available build/configuration state only. It does not contact deployed endpoints, external identity providers, registries, brokers, cloud APIs, or other shared services.
|
||||
|
||||
Launch several `research` agents in parallel. They return structured facts to the parent and do not write files.
|
||||
|
||||
**Agent 1a: Product, stack, and local operation**
|
||||
|
||||
```text
|
||||
Read the target at <target>. Do not use network access. Return:
|
||||
1. Product type, users, operators, and ordinary trust-sensitive actions.
|
||||
2. Languages, frameworks, build system, runtimes, and locally visible deployment models.
|
||||
3. Repository-relative entry points and subsystem boundaries.
|
||||
4. Exact build and test commands that could run offline with local dependencies, their expected write locations, and the target-controlled inputs they process. Do not run them during reconnaissance.
|
||||
5. Comparable software or protocol visible from local documentation and dependencies. If no useful comparison is source-grounded, say so.
|
||||
6. Missing local toolchains or runtime facts that limit bounded execution.
|
||||
Return only source facts with repository-relative file:line references.
|
||||
```
|
||||
|
||||
**Agent 1b: Principals, authority, and controls**
|
||||
|
||||
```text
|
||||
Read all source that establishes identity, authorization, isolation, and privilege. Map:
|
||||
1. Each lower-trust principal and the actions it has by design.
|
||||
2. Authentication or peer identity at each entry surface.
|
||||
3. Per-resource authorization and tenant/owner scope.
|
||||
4. Process, browser, workload, CI, plugin, model/tool, device, or local-IPC authority.
|
||||
5. Privilege changes, confirmation, revocation, recovery, and fallback paths.
|
||||
6. Which controls are source-visible and which depend on an unobserved deployment fact.
|
||||
Return trust boundaries and control locations with repository-relative file:line references. Do not infer live reachability.
|
||||
```
|
||||
|
||||
**Agent 1c: Entry surfaces, copies, and sinks**
|
||||
|
||||
```text
|
||||
Inventory every source-visible place external or lower-trust input enters:
|
||||
- HTTP/browser, RPC/message/protocol, files/archive/document, CLI/env/config, plugins/dependencies/CI, cloud events/IAM selectors, model context/tool arguments, mobile/deep-link/webview, and local IPC.
|
||||
For each surface, follow major transformations, stored or derived copies, and security-relevant sinks. Record source-visible limits and parallel paths to the same effect.
|
||||
Return repository-relative paths and line numbers. Be complete, but do not execute or send inputs.
|
||||
```
|
||||
|
||||
**Agent 1d: Local execution and deployment visibility**
|
||||
|
||||
```text
|
||||
Read tests, build definitions, manifests, packaging, and maintained environment overlays. Return:
|
||||
1. Small offline tests or existing fixtures that could validate trust boundaries with dummy data inside the required OS-enforced sandbox.
|
||||
2. Processes that could use an isolated loopback network namespace without external or shared dependencies.
|
||||
3. Commands that would fetch dependencies, publish artifacts, contact paid/provider APIs, or affect shared state; mark them prohibited for this run.
|
||||
4. Deployed controls and attachments that source cannot establish and therefore require needs_validation if decisive.
|
||||
5. The final active source path for each deployment mode only where the repository selects it deterministically.
|
||||
6. Whether the local platform can enforce an empty allowlisted environment, no external network, read-only target/toolchain mounts, scratch-only writes, and explicit CPU, memory, process, file-size, disk, and wall-clock limits. Missing controls block target-controlled execution.
|
||||
7. Whether trusted parent-side code can promote predeclared scratch files with path-confined no-follow descriptor traversal, nonblocking regular-file checks, no-follow traversal of every destination parent, exclusive regular-file destination creation, and explicit per-file and cumulative size bounds. Missing promotion controls block use of scratch files as evidence.
|
||||
```
|
||||
|
||||
Add focused reconnaissance agents for materially distinct deployment modes or subsystems that these four do not map. Do not silently omit them: if the budget gate in `SKILL.md` blocks a focused agent, launch nothing for it, seed the unmapped area as a `deferred` ledger unit with reason `budget_cannot_reserve_critics_and_validation`, and disclose the gap in the report.
|
||||
|
||||
## Prior-run input
|
||||
|
||||
Before selecting work, the parent reads every available prior `coverage-ledger.json` and `findings.json` for the same repo:
|
||||
|
||||
- Compare the source locations, controls, conditions, and source-derived identity for every prior record and unit against the current source.
|
||||
- Carry an unchanged prior `confirmed` record into the current candidate set, with the same fingerprint, only when its relevant source, conditions, and qualifying evidence still apply. Link it to a current `planned` unit with `prior_status: "prior_confirmed_same_source"` and put only that root cause on the hunter exclusion list. The Phase 3 verifier that re-verifies the carried record becomes that unit's assignment owner; its source re-check is the unit's first check and moves the unit to `candidate` with the carried fingerprint.
|
||||
- Build a current planned `prior_confirmed_changed_source` revalidation unit when any relevant source or condition changed. Do not exclude that root cause from hunting or assume the prior verdict still applies.
|
||||
- Build current work units for every prior `needs_validation`, `deferred`, `blocked`, `out_of_scope`, and changed-source unit. These states are priority input, never deduplication or suppression keys.
|
||||
- Carry a still-blocked prior `needs_validation` record with the same fingerprint only after current source supports its trace. Link it to a current `planned` unit with `prior_status: "prior_needs_validation"`; the record keeps the unresolved blocker. The Phase 3 verifier that re-checks the carried record becomes that unit's assignment owner; its re-check is the unit's first check and moves the unit to `candidate` with the carried fingerprint. Include the record in final verification.
|
||||
- Treat prior rejected records as stale claims unless current evidence changes the failed trace or missing condition. An unchanged rejection suppresses only that exact claim, not review of the coverage unit.
|
||||
- Record missing or incompatible ledgers instead of treating them as empty coverage.
|
||||
|
||||
State paths and source refs used in `run-metadata.json`. Summarize only the coverage consequences in `architecture.md`.
|
||||
|
||||
## Architecture summary and companion selection
|
||||
|
||||
The parent synthesizes `<output-dir>/architecture.md`, with a hard cap of about 1,000 words. Include:
|
||||
|
||||
1. Product, principals, normal authority, and protected resources.
|
||||
2. The comparable-software baseline from Agent 1a, when one is source-grounded: what security trade-offs the comparable accepts. Use it to calibrate effort and severity, never to dismiss a demonstrated finding; if the comparable shares a defect pattern that has mattered in practice, that strengthens the finding. Omit this line when no meaningful comparable exists.
|
||||
3. Tech stack, source-visible deployment paths, and offline build/test limits.
|
||||
4. Entry surfaces and the important source-to-sink or lifecycle paths.
|
||||
5. Trust boundaries and the strongest source-visible control on each.
|
||||
6. Repository-relative starting paths.
|
||||
7. Prior coverage gaps, changed-source and blocked revalidation targets, and same-source confirmed exclusions.
|
||||
8. A short companion-selection summary derived from [ATTACK-CLASSES.md](ATTACK-CLASSES.md): selected files and the source-visible boundaries that require them.
|
||||
|
||||
Keep the assignment-level ordinary block, selected companion blocks, and excluded blocks with reasons in each ledger unit, not in `architecture.md`. This keeps the architecture cap valid for large runs and makes the exact hunter prompt map machine-checkable.
|
||||
|
||||
Do not select a companion file merely because the language or dependency name appears. Select it because reconnaissance found the trust-sensitive boundary described by its `When to use this file` section. Do not exclude a visible boundary just because another agent will review a related class.
|
||||
|
||||
## Deterministic coverage ledger
|
||||
|
||||
The parent writes `<output-dir>/coverage-ledger.json` as a top-level JSON array. Derive one unit for every material combination of entry surface, trust boundary, subsystem, and applicable ordinary or companion attack class at the granularity the run profile sets (`quick` uses one all-in-scope subsystem identity; `deep` adds lifecycle modes). For a scoped run, seed in-scope surfaces for assignment and retain discovered excluded surfaces as `out_of_scope` units so later full runs can turn them into current work.
|
||||
|
||||
Each dimension has a human label and a stable source-derived value in `canonical_refs`. Use the same canonical reference for the same source object across runs even if its display label changes. Suitable references include a repository-relative entry path plus exported scope, a route or message identity defined in source, the source control that defines a boundary, a repository package path, and the exact attack-class block reference. A block reference is `FILE.md#` plus the exact class name as written in bold or as a heading in that file — a stable identifier matched against the file text, not a rendered HTML anchor. For companion section blocks, use the heading text before any parenthetical qualifier (for example `Core discipline`). Do not derive references by lowercasing or slugging display labels.
|
||||
|
||||
Derive `coverage_id` without lossy slugs:
|
||||
|
||||
1. Require every reference to be Unicode NFC with valid scalar values, visible content, no control, format, line/paragraph separator, or default-ignorable code point, and no surrounding whitespace.
|
||||
2. Encode its UTF-8 bytes with RFC 3986 percent encoding: leave only `A-Z a-z 0-9 - . _ ~` unescaped and use uppercase `%HH` for every other byte.
|
||||
3. Join encoded `surface`, `boundary`, `subsystem`, and `attack_class` references with `::`; append encoded `lifecycle` when present.
|
||||
|
||||
Use the fixed canonical value `profile/quick/all-in-scope-subsystems` for the quick profile's coarsened subsystem dimension. Do not include wave number, agent, verdict, severity, or line number in a reference or ID. Sort units lexicographically by `coverage_id` before each assignment. Fail on every duplicate ID. If duplicate IDs have different semantic fields, treat that as a canonical identity collision; never merge or silently overwrite them. The validator also rejects one semantic tuple represented by different canonical references.
|
||||
|
||||
Each unit records:
|
||||
|
||||
```json
|
||||
{
|
||||
"coverage_id": "...",
|
||||
"canonical_refs": {
|
||||
"surface": "src/router.ts#POST /users/:id",
|
||||
"boundary": "src/authz.ts#requireOwner",
|
||||
"subsystem": "packages/api",
|
||||
"attack_class": "ATTACK-CLASSES.md#Access control"
|
||||
},
|
||||
"surface": "...",
|
||||
"boundary": "...",
|
||||
"subsystem": "...",
|
||||
"attack_class": "...",
|
||||
"starting_paths": ["repo/relative/path"],
|
||||
"ordinary_attack_class_block": "ATTACK-CLASSES.md#Access control",
|
||||
"selected_companion_blocks": ["FILE.md#section"],
|
||||
"excluded_blocks": [{"block": "FILE.md#section", "reason": "..."}],
|
||||
"prior_status": "new|prior_confirmed_same_source|prior_confirmed_changed_source|prior_needs_validation|prior_deferred|prior_blocked|prior_out_of_scope|prior_covered_same_source|prior_covered_changed_source|prior_rejected_claim_changed|none",
|
||||
"attempts": [],
|
||||
"wave": 1,
|
||||
"status": "planned",
|
||||
"agent_id": null,
|
||||
"reviewed_paths": [],
|
||||
"local_checks": [],
|
||||
"result_fingerprints": [],
|
||||
"unresolved": []
|
||||
}
|
||||
```
|
||||
|
||||
When `lifecycle` is material, add both `canonical_refs.lifecycle` and a human `lifecycle` field. `ordinary_attack_class_block` is null only when no ordinary block applies. The selected companion list includes each applicable class plus its companion `Core discipline`, `Universal moves`, and `Validation rules`; `excluded_blocks` records every considered but unselected block and the source fact that excludes it.
|
||||
|
||||
The parent may add bookkeeping fields but keeps the semantic fields above stable. In `prior_status`, `new` marks a surface first seen in this run when compatible prior ledgers exist; `none` marks a unit seeded when no compatible prior ledger is available. Prior `deferred`, `blocked`, `out_of_scope`, and changed-source units initialize as current `planned` work when now in scope. A prior same-source covered unit remains visible in the current ledger; assign changed source, important lifecycle paths, and exact conflicts first, then use the coverage critic to decide whether it needs another pass.
|
||||
|
||||
`attempts` is an append-only archive for evidence-bearing assignments that a coverage critic reopens. Before reassignment, append the prior unit's exact `wave`, `status`, `agent_id`, `reviewed_paths`, `local_checks`, `result_fingerprints`, and `unresolved`, plus the critic's source-backed `reassignment_reason`. Only `blocked`, `covered`, and `candidate` states can be archived. Archived attempts retain the same state and evidence invariants as live units, use strictly increasing waves below the current wave, and retain their producing owners and artifacts. The next assignment increments `wave`, uses a fresh owner, and starts with empty live evidence. If the profile or budget prevents another assignment, increment `wave` and use live `deferred` state with null owner, empty evidence, and the stop reason. Never copy an archived owner's checks or artifacts into the live state. A later live terminal state contains only the new attempt's evidence; the archive remains unchanged.
|
||||
|
||||
Enforce this state table exactly:
|
||||
|
||||
| Status | Unit `agent_id` | `reviewed_paths` / `local_checks` | `result_fingerprints` | `unresolved` |
|
||||
|---|---|---|---|---|
|
||||
| `planned` | null | empty | empty | empty |
|
||||
| `not_applicable`, `out_of_scope`, `deferred` | null | empty | empty | nonempty reason |
|
||||
| `in_progress` | canonical owner | empty | empty | empty |
|
||||
| `blocked` | canonical owner | both nonempty owned partial evidence | empty | nonempty blocker |
|
||||
| `covered` | canonical owner | both nonempty | empty | empty |
|
||||
| `candidate` | canonical owner | both nonempty | nonempty | optional |
|
||||
|
||||
Canonical agent IDs match `^[a-z0-9][a-z0-9_-]{0,63}$` and are not Windows device names. Lowercase is mandatory, so one ledger cannot contain case-fold aliases. The unit `agent_id` records the assignment owner. Every check records its own `agent_id` and nonempty `reviewed_paths`; the unit-level `reviewed_paths` is exactly their union. A source-only check uses `artifact: null`. A local check requires a regular file promoted only by trusted parent-side code under exactly `agents/<check.agent_id>/artifacts/`. This lets hunter and verifier checks coexist in one unit. Scratch paths, output-root files, symlinks, special files, and another check owner's artifacts are not evidence.
|
||||
|
||||
The ledger is the coverage claim. An architecture summary, agent count, or generic "auth reviewed" sentence is not coverage evidence. Phase 2 closes units only from the paths and checks in a hunter's structured result.
|
||||
|
||||
Run `node <skill-dir>/validate-coverage-ledger.cjs <output-dir>/coverage-ledger.json` after seeding, after every parent update, and before Phase 6. The validator rejects input beyond 5 MiB, 64 nesting levels, 10,000 units, 1,000 entries in a nested collection, or 500,000 traversed values, and caps reported validation errors at 100. In practice the 5 MiB byte limit holds roughly 2,000-5,000 realistic units, so it binds before the 10,000-unit cap. Fix every error before assigning work or making a coverage claim.
|
||||
@@ -0,0 +1,78 @@
|
||||
# Resource Exhaustion and Availability Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when untrusted requests, messages, files, tenant state, or agent work can consume CPU, memory, disk, connections, worker slots, paid APIs, or queue capacity, or can deadlock/crash a shared service. This domain distinguishes a source-reviewable availability vulnerability from a general performance issue. Never validate by stressing a shared or live service.
|
||||
|
||||
Use `MEMORY-SAFETY-AND-BINARY.md` for memory-integrity defects and `PROTOCOLS-RPC-AND-MESSAGING.md` for broker delivery logic. A reachable fatal error belongs here for shared impact even when the underlying parser is covered elsewhere.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- Require an input-to-cost path, a missing effective bound, and impact on another user, shared service, safety function, or operator-owned spend. Self-limiting work in the requester's own process is not a service vulnerability.
|
||||
- A missing rate limit is not enough. Check body/message/file caps, concurrency, queues, deadlines, database constraints, upstream gateways, and per-tenant quotas before calling a path unbounded.
|
||||
- Do not run stress, saturation, or production tests. Use asymptotic analysis, small boundary fixtures, mocked paid calls, strict local resource limits, and deterministic cancellation tests.
|
||||
- State attacker cost, service work, persistence, scope, and recovery. One bounded input with superlinear or persistent shared effect is materially different from sustained volume.
|
||||
- Use `confirmed` for source-visible bounds failures demonstrated safely. Use `needs_validation` when upstream caps, deployed topology, autoscaling, paid quota, or recovery behavior is outside the repository.
|
||||
```
|
||||
|
||||
## Computational amplification attack classes (subagent_type: `general`)
|
||||
|
||||
**Superlinear parsing, matching, or evaluation**
|
||||
Small accepted input drives catastrophic regex backtracking, nested parsing, recursive validation, symbolic evaluation, graph traversal, template expansion, or adversarial sort/hash behavior. Derive accepted depth/cardinality and complexity, then demonstrate a bounded growth curve locally.
|
||||
|
||||
**Decompression and representation amplification**
|
||||
Compressed, sparse, nested, aliased, or encoded input expands far beyond the checked transfer or file size. Verify limits after every expansion and across parser stages, including archives, images, fonts, structured documents, and protocol compression tables.
|
||||
|
||||
**Database and downstream query amplification**
|
||||
A small request creates broad scans, pathological joins, fan-out, unbounded sort/aggregation, or many downstream calls because query depth, filter cardinality, pagination, or expansion fields are not bounded. Confirm authorization does not intentionally permit the same resource scope.
|
||||
|
||||
## Resource accumulation attack classes (subagent_type: `general`)
|
||||
|
||||
**Unbounded buffering and cardinality**
|
||||
Bodies, out-of-order streams, uploads, sessions, unique cache keys, metrics labels, log fields, subscriptions, or pending jobs accumulate without per-item and aggregate limits. Find cleanup and expiration on disconnect, timeout, cancellation, and partial parse.
|
||||
|
||||
**File descriptor, handle, and temporary-resource leaks**
|
||||
Malformed or canceled work misses cleanup and retains sockets, files, database cursors, timers, subprocesses, temporary files, or object references. Confirm the leak repeats through bounded local iterations and affects a shared pool.
|
||||
|
||||
**Detached work after cancellation**
|
||||
Client timeout, disconnect, canceled job, or failed authorization returns control but leaves database, model, network, or worker work running. Trace cancellation and deadline propagation through every layer.
|
||||
|
||||
## Quota and scheduling attack classes (subagent_type: `general`)
|
||||
|
||||
**Pre-authentication work imbalance**
|
||||
Expensive parsing, key lookup, cryptography, decompression, or external requests happen before authentication and the earliest size/rate gate. Compare minimal requester effort to shared service cost and check upstream limits.
|
||||
|
||||
**Quota-accounting scope and reset gaps**
|
||||
Accounting uses attacker-influenceable IP, route, tenant, key prefix, task ID, or other dimension, allowing one principal's work to escape its intended budget or consume another principal's allocation. Review integer overflow, distributed races, retries, reconnects, and account switching.
|
||||
|
||||
**Worker, pool, and priority starvation**
|
||||
Low-priority or attacker-controlled jobs hold shared locks, workers, database pools, event-loop turns, or scheduler priority needed by unrelated users. Require a path that bypasses queue/concurrency fairness or retains a slot beyond its deadline.
|
||||
|
||||
## Failure and recovery attack classes (subagent_type: `general`)
|
||||
|
||||
**Reachable fatal error or deadlock**
|
||||
An untrusted input reaches `panic`, abort, fatal assertion, unhandled exception, process exit, lock cycle, or infinite loop in a shared process. Confirm supervisor scope and whether one worker or the whole service becomes unavailable. A restarted isolated worker may reduce impact but does not erase the defect.
|
||||
|
||||
**Retry storm and fail-open amplification**
|
||||
Timeouts, dependency errors, partially processed messages, or health-check failures trigger synchronized or unbounded retries without jitter, ceilings, circuit breaking, or deduplication. Verify one bounded failure source can create persistent aggregate work.
|
||||
|
||||
**Poison-record and head-of-line blocking**
|
||||
One malformed record or message repeatedly fails at the front of a shared queue, partition, startup scan, migration, or recovery loop. Review skip/quarantine policy, offsets, and whether other tenants share the blocked unit.
|
||||
|
||||
**Unsafe recovery and capacity rollback**
|
||||
A restart, restore, fallback, or cleanup path rebuilds unbounded state, ignores current quotas, or restores the input that immediately repeats failure. Recovery correctness is part of availability.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Build an input-to-resource table: earliest accepted size/cardinality, work before auth, downstream fan-out, persistence, shared pool, limit and cleanup owner, recovery.
|
||||
- Compare aggregate limits with per-object limits. Ten thousand valid one-byte items may evade a per-message cap while exhausting tenant-wide or process-wide state.
|
||||
- Validate only in an isolated fixture with strict CPU/memory/time limits and small growth points. Mock external and paid calls and stop once the missing bound or cancellation is observable.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Name untrusted input, requester work, service amplification or retained resource, shared blast radius, and recovery. Missing limits without concrete shared impact are hardening.
|
||||
2. Confirm no source-visible upstream, parser, queue, tenant, or framework bound prevents the path. Unknown deployed controls require `needs_validation`.
|
||||
3. For superlinear behavior, establish the accepted complexity and bounded local growth. For leaks, show repeatable retention after cleanup should occur. For fatal paths, identify process/supervisor isolation.
|
||||
4. Prioritize by low requester work, unauthenticated reachability, cross-tenant scope, persistence, and poor recovery; do not validate with availability impact.
|
||||
5. Return `confirmed` only with safe local proof and meaningful shared effect. Return `needs_validation` with the exact upstream limit, topology, quota, or recovery observation an owner must check.
|
||||
192
.agents/skills/security-audit/SKILL.md
Normal file
192
.agents/skills/security-audit/SKILL.md
Normal file
@@ -0,0 +1,192 @@
|
||||
---
|
||||
name: security-audit
|
||||
description: Security guidance and vulnerability review for codebases, APIs, services, CLI tools, libraries, and daemons. Use for security questions, focused reviews, vulnerability research, security audits, or pen tests. Run the complete workflow only for explicit codebase audit or pen-test requests, full/comprehensive/end-to-end reviews, or requested report artifacts.
|
||||
---
|
||||
|
||||
# Security Audit
|
||||
|
||||
Find vulnerabilities that violate a real trust boundary, then give owners the source evidence, safe reproduction, priority, and smallest effective fix. This is a defensive, source-first workflow. A candidate without a concrete affected principal, resource, or security outcome is not a confirmed finding.
|
||||
|
||||
## Operating modes
|
||||
|
||||
This skill is guidance by default. Loading it does not authorize the complete audit workflow or file creation.
|
||||
|
||||
- **Guidance mode**: For security questions, focused reviews, methodology, triage, or investigation of specific findings, use only the relevant parts of this skill. Do not automatically run all six phases, create an output directory, or write audit artifacts. You may launch focused agents when useful; they return results to the current task.
|
||||
- **Full audit mode**: Use the complete workflow when the user explicitly asks to audit or pen-test a codebase, asks for a full, comprehensive, or end-to-end security review, or requests report artifacts. Run all six phases and write the files defined below.
|
||||
|
||||
If the request could mean either mode, ask one focused question before creating files or starting the complete workflow.
|
||||
|
||||
## Platform terminology
|
||||
|
||||
This skill is agent-neutral:
|
||||
|
||||
- **Parent** is the agent that coordinates the run and owns shared state.
|
||||
- **Task tool** is the platform's delegation or sub-agent mechanism.
|
||||
- **`research` agent** is a delegated agent for focused source exploration and factual verification.
|
||||
- **`general` agent** is a delegated agent for broad investigation and bounded local execution.
|
||||
- **`subagent_type:`** in a heading names which of these two delegated agent roles runs that work.
|
||||
|
||||
Use equivalent platform capabilities while preserving role, write-isolation, prompt, and independence boundaries.
|
||||
|
||||
## Universal execution safety
|
||||
|
||||
These rules apply in both operating modes. Source inspection is read-only. Run target-controlled builds, tests, processes, browsers, emulators, fuzzers, and fixture processing only inside an OS-enforced sandbox that provides all of these controls:
|
||||
|
||||
- no external network; use only an isolated loopback namespace when the check needs local client/server traffic;
|
||||
- an empty environment populated from an explicit allowlist with safe values, with scratch-local `HOME`, temporary directories, and caches;
|
||||
- a read-only target and toolchain, with the target-controlled process able to write only inside its assigned `scratch/` directory; and
|
||||
- explicit low CPU, memory, process, file-size, disk, and wall-clock limits.
|
||||
|
||||
The agent, outside the target-controlled process, may make a disposable source copy in an assigned `scratch/` directory when a build must write beside source. In guidance mode, do not retain target-controlled files. In full audit mode, only trusted parent-side code may promote the minimum non-secret result to retained `artifacts/` using the procedure under Write isolation. Never expose a retained output directory (other than the agent's own assigned `scratch/`), another agent's directory, the host home directory, credentials, sockets, or shared services to target code. Do not install dependencies or let builds fetch them. Use only tools and dependencies already available locally. If every control cannot be enforced, do not execute target code: report the missing sandbox capability as a needs-validation blocker and give a safe validation plan.
|
||||
|
||||
Use dummy principals, fixtures, and secrets. Do not probe deployed endpoints, external services, shared infrastructure, production identities, other users' data, or live control planes. Do not test availability against a live or shared process, publish artifacts, alter releases, spend paid API quota, or continue beyond the minimum local effect needed to establish a defect. If the decisive fact is outside source or the sandboxed fixture, report it as needing validation.
|
||||
|
||||
## Full audit setup
|
||||
|
||||
In full audit mode, resolve these values before reconnaissance:
|
||||
|
||||
- **Skill directory**: the absolute directory containing this `SKILL.md`.
|
||||
- **Target**: the absolute repository root under review.
|
||||
- **Repo name**: a stable repository identifier from the directory or local Git remote.
|
||||
- **Output directory**: a new writable directory outside the target, defaulting to `~/security-audit-skill/<repo-name>/run-<N>`, where `<N>` is the next unused integer. Use a directory inside the target only when the user explicitly selects it and the parent verifies that version control ignores the whole directory. Otherwise stop and request an external path.
|
||||
- **Source ref**: the reviewed commit and whether the worktree is dirty. Do not treat unreviewed generated or modified files as another revision.
|
||||
|
||||
### Write isolation
|
||||
|
||||
The parent creates and is the only writer of shared run files:
|
||||
|
||||
- `run-metadata.json`
|
||||
- `architecture.md`
|
||||
- `coverage-ledger.json`
|
||||
- `findings.json`
|
||||
- `REPORT.md`
|
||||
- `FINDINGS-DETAIL.md`
|
||||
- `NEEDS-VALIDATION.md`
|
||||
|
||||
Each hunter or verifier receives a unique root under `<output-dir>/agents/<agent-id>/`, with separate `scratch/` and `artifacts/` directories. Canonical agent IDs match `^[a-z0-9][a-z0-9_-]{0,63}$` and must not equal a Windows device name such as `con`, `prn`, `aux`, `nul`, `com1` through `com9`, or `lpt1` through `lpt9`. Lowercase IDs prevent case-fold collisions. The agent and every target-controlled process may write only to `scratch/`; retained `artifacts/` is parent-owned, is never exposed to the sandbox, and is writable only by trusted parent-side promotion code. Agents may not change shared files, target source, retained artifacts, or another agent's directory. Do not use `/tmp` or the host home directory as a writable fallback.
|
||||
|
||||
Before execution, the parent opens and retains trusted, non-inheritable directory descriptors for the agent's `scratch/` and `artifacts/` roots, and records an allowlist of expected scratch-relative artifact files plus explicit per-file and cumulative byte limits. Never pass those descriptors to the agent or sandbox. After the sandbox and all its processes terminate, trusted parent-side code promotes each allowlisted file separately:
|
||||
|
||||
1. Validate the declared relative path: reject absolute, empty, `.`, `..`, or symlinked components.
|
||||
2. Walk each parent component from the retained scratch-root descriptor with no-follow directory-relative operations; never reopen by path.
|
||||
3. Open the leaf no-follow and nonblocking.
|
||||
4. Verify with `fstat` that it is a regular file with link count exactly one and within the recorded per-file and cumulative byte limits.
|
||||
5. Enforce those limits again while reading from that descriptor.
|
||||
6. Copy exactly the verified size, repeat `fstat`, and reject a changed identity, type, link count, or size.
|
||||
7. For the destination, walk every parent component from the retained artifacts-root descriptor with no-follow directory-relative operations; require each existing component to be a real directory, and create any missing directory exclusively before reopening and verifying it no-follow.
|
||||
8. Create the leaf exclusively without following links, verify that the opened destination is a regular file with link count exactly one, and copy from the verified source descriptor without reopening either path.
|
||||
9. Use equivalent race-safe APIs on non-POSIX systems.
|
||||
10. Never recursively copy or glob scratch, extract an archive into artifacts, or open or promote a symlink, FIFO, socket, device, directory, hard-linked file, changing file, or file that exceeds its bound.
|
||||
11. If any check is unavailable, cannot be enforced, or fails, discard the scratch entry; if it is decisive evidence, retain `needs_validation` with the exact promotion blocker.
|
||||
|
||||
[HUNTING.md](HUNTING.md) and [VALIDATION-AND-REPORTING.md](VALIDATION-AND-REPORTING.md) carry this procedure as one identical fenced block for hunter and verifier prompts; it states the same rules in the same order as this list.
|
||||
|
||||
For a reproduced check, record the command, exact test input, sandbox limits, and only the allowlisted environment variable names plus safe non-secret values needed to reproduce it. Never capture or copy the ambient environment, inherited variables, credential values, authentication state, or unrelated host paths. Launch from an empty environment rather than trying to redact one after execution.
|
||||
|
||||
Before delegation, the parent writes `run-metadata.json` with at least `run_id`, `repo`, `target`, `source_ref`, `profile`, `scope_paths`, `budget` (null if unset), `execution_policy: "sandboxed-source-and-local-only"`, selected companion files, prior-run paths, shared-file owners, and `run_status: "in_progress"`. Update metadata only when those facts change; candidate state belongs in the coverage ledger and `findings.json`.
|
||||
|
||||
## Full audit planning
|
||||
|
||||
The coverage, prior-run, profile, and budget requirements in this section apply only in full audit mode.
|
||||
|
||||
### Coverage and prior runs
|
||||
|
||||
No one pass is complete. Build a deterministic coverage plan before hunting and update it after every agent result. [RECONNAISSANCE.md](RECONNAISSANCE.md) defines the stable coverage units and [HUNTING.md](HUNTING.md) defines coverage-critic waves. The parent alone updates the ledger.
|
||||
|
||||
If prior runs exist, read every compatible `coverage-ledger.json` and `findings.json` before planning the current run:
|
||||
|
||||
1. Compare the relevant current source with each prior record and unit. A prior source ref alone is not evidence that a path is unchanged.
|
||||
2. Carry a prior `confirmed` record into the current candidate set only when its relevant source and conditions are unchanged and its evidence still meets the current contract. Link it to a current ledger unit seeded `planned`, preserve its fingerprint, exclude only that carried root cause from hunters, and send the carried record through the current final verification path; the Phase 3 verifier that re-checks it becomes that unit's assignment owner and moves it to `candidate`.
|
||||
3. When relevant source for a prior `confirmed` record changed, create a current planned revalidation unit. Do not put that record on the hunter exclusion list. It remains confirmed only if current independent validation establishes the current path and result.
|
||||
4. Make prior `needs_validation`, `deferred`, `blocked`, `out_of_scope`, and any changed-source unit current work. A still-external `needs_validation` record may be carried only after the current source trace is checked and linked by fingerprint to a current `planned` unit whose verifier re-check supplies its owner and evidence; the record keeps the unresolved blocker. These prior states never suppress a current unit.
|
||||
5. A prior same-source covered unit may inform priority, but it remains visible in the current ledger. A prior `rejected` record suppresses only the unchanged failed claim, not coverage of its unit; changed evidence creates current work.
|
||||
6. Read the prior profile and scope. A prior `quick` or scoped ledger contributes only its recorded evidence and gaps, never an implied "rest is fine."
|
||||
|
||||
If no prior ledger exists, say so in the final coverage statement. Never imply that one run exhausts the target.
|
||||
|
||||
### Run profiles and scope
|
||||
|
||||
During full audit setup, pick a profile from the user's request or propose one from the target's size and stakes. Record it in `run-metadata.json` (`profile`, `scope_paths`) and state it in the report. The default is `standard`.
|
||||
|
||||
- **`quick`** — a bounded pass for small targets, re-runs, or a fast first look. Coarsen ledger units to surface × boundary × attack class (subsystem uses the fixed canonical `profile/quick/all-in-scope-subsystems` identifier), run exactly one hunter wave followed by exactly one final coverage-critic pass, and use one fresh verifier per candidate for both candidate validation and final record verification. Do not launch a follow-up hunter wave: record the critic's accepted discoveries and reassignments as `deferred`.
|
||||
- **`standard`** — the workflow as written.
|
||||
- **`deep`** — for high-stakes or large targets. Split ledger units per subsystem and lifecycle mode, run critic waves to a clean pass, keep candidate validation and final record verification as separate fresh agents, and give `prior_covered_same_source` units an independent second pass.
|
||||
|
||||
A **scoped run** audits a subset: named paths, one subsystem, one companion domain, or the diff between two source refs. Seed ledger units only for in-scope surfaces and record everything else as `out_of_scope` — never as `covered`. A scoped or `quick` run must present itself as partial coverage.
|
||||
|
||||
Profiles change breadth and redundancy, never the evidence bar. Do not scale away the candidate gate, the source/local execution boundary, `needs_validation` discipline, schema validation, or independent verification of `confirmed` records.
|
||||
|
||||
#### Cost budget
|
||||
|
||||
The ledger makes spend countable: one unit is roughly one hunter assignment, and one surviving candidate is one or two verifier assignments depending on profile. When the user sets a budget — or the parent proposes one for a large target — record `budget` in `run-metadata.json` as a maximum number of agent invocations across all phases.
|
||||
|
||||
Apply the strict budget gate before launching any reconnaissance agent. Reserve the four baseline reconnaissance calls, one final post-wave critic for `quick` or one post-wave plus one distinct final-clean critic for `standard`/`deep`, and at least one verifier call. Add focused reconnaissance only after repeating this gate for each extra call. If the requested budget cannot fund that minimum, launch no agent: ask for a larger budget, narrower scope, or different profile. If the request remains unchanged, set `run_status: "incomplete"` with `incomplete_reason: "budget_cannot_fund_reconnaissance_and_reserves"` and report that no audit pass ran.
|
||||
|
||||
Spend it in this order:
|
||||
|
||||
1. Count reconnaissance, every post-wave critic, and the separate final-clean critic as agent invocations.
|
||||
2. **Reserve critics and validation before hunting.** For `quick`, reserve its one post-wave final critic. Before every `standard` or `deep` hunter wave, reserve one immediate post-wave critic plus one distinct final-clean critic. Also reserve verifier cost from the profile (about 1 or 2 agents per expected candidate; when in doubt reserve 30% of the balance after critic reservation). Never assign hunters into either reserve.
|
||||
3. Assign hunters to units in priority order until the hunting allowance is spent. Spend the reserved post-wave critic immediately after that wave; keep the final-clean and validation reserves intact.
|
||||
4. Before a later wave, reserve its new post-wave critic again. If the remaining budget cannot cover the required critic calls and validation reserve, launch no hunters from that wave, mark its planned units `deferred` with reason `budget_cannot_reserve_critics_and_validation`, and use the retained final-clean critic to record the resulting gap.
|
||||
|
||||
Before wave 1, update the pre-recon estimate with seeded units, implied hunter count, mandatory critic calls, validation reserve, and whether the remaining budget covers the plan. If it clearly cannot, say so and propose either a tighter scope or a coarser profile instead of silently thinning evidence. If later facts consume the required final-critic reserve, launch no hunters, mark all planned work deferred, set the run incomplete with reason `critic_budget_exhausted`, and make no complete-coverage claim.
|
||||
|
||||
A strict total-agent budget can still be exceeded by an unexpectedly large candidate set or by a material Phase 5 replacement that needs another independent verifier. If the remaining budget cannot validate every candidate, stop hunting, validate candidates in fingerprint order while the budget permits, and set `run_status: "incomplete"` plus `incomplete_reason: "validation_budget_exhausted"`. Keep each unvalidated fingerprint linked to a `candidate` ledger unit with that unresolved reason. Do not put an unvalidated candidate in `findings.json`, relabel it `needs_validation`, or report the run as complete. Phase 6 may produce a partial report only if its first section states that candidate validation is incomplete and lists the affected fingerprints and units. Never exceed a user-set strict budget silently.
|
||||
|
||||
## Core principles
|
||||
|
||||
### Require a boundary and result
|
||||
|
||||
For every candidate, name the lower-trust principal, accepted input or action, intended control, crossed boundary, affected principal or resource, and concrete observed or owner-observable result. Do not elevate a missing best practice, guessed deployment behavior, generic parser crash, or self-impact into a security finding.
|
||||
|
||||
### Use bounded local evidence
|
||||
|
||||
Static analysis establishes the source path. Sandboxed local tests resolve behavior when all execution controls are available: a minimal function harness, existing unit test, small parser fixture, dummy-tenant integration test, locally rendered configuration, or bounded isolated-loopback client. Stop at a wrong return value, unauthorized dummy record, sanitizer finding, policy difference, or other minimum effect. Do not extend the local check beyond the minimum boundary result or produce persistence, post-fault, or concealment material.
|
||||
|
||||
### Respect source visibility
|
||||
|
||||
Deployment controls, proxy behavior, provider settings, browser headers, identity policy, broker ACLs, packaging, and topology are real controls. If they are required and absent from the repository, do not assume either presence or absence. Use `needs_validation` with the exact missing fact and a safe owner-observed or local plan.
|
||||
|
||||
### Separate priority from certainty
|
||||
|
||||
Only `confirmed` records receive severity. Likelihood and impact must reflect the demonstrated conditions and result; overall severity cannot exceed demonstrated impact. `needs_validation` means a specific source-grounded boundary hypothesis is blocked, not a low-confidence confirmed vulnerability, and it has no severity.
|
||||
|
||||
Calibrate overall severity with these anchors:
|
||||
|
||||
- **critical** — an unauthenticated actor gains code execution, full data-store access, or takeover of arbitrary accounts.
|
||||
- **high** — an actor fully defeats an explicit security control with real consequences: authentication bypass, cross-tenant read or write, stored script execution affecting other users, authenticated code execution, or an unauthenticated remote stop of a shared service.
|
||||
- **medium** — a real boundary violation with limited blast radius, uncommon preconditions, or consequences confined to a narrow resource set.
|
||||
- **low** — disclosure of non-secret internals, or an effect requiring sustained effort for minimal gain.
|
||||
- **informational** — a confirmed but minimal-impact observation, useful mainly as a prerequisite inside a larger finding.
|
||||
|
||||
The high/medium discriminator: does the demonstrated result fully defeat an explicit control for an action with real consequences, or only weaken it? If you cannot state the concrete damage, the severity is lower than it feels.
|
||||
|
||||
### Recommend the smallest effective source fix
|
||||
|
||||
For each confirmed finding, identify the invariant the code must enforce and the narrowest source change that enforces it at the last trusted decision point. Prefer specific repository-relative changes and regression tests over generic hardening advice. The audit describes fixes; it does not modify target source.
|
||||
|
||||
## Full audit workflow
|
||||
|
||||
In full audit mode, follow all six phases in order:
|
||||
|
||||
1. **Reconnaissance** — map the source, trust boundaries, local build paths, companion selections, prior evidence, and initial deterministic coverage ledger with [RECONNAISSANCE.md](RECONNAISSANCE.md).
|
||||
2. **Coverage-led hunting waves** — assign isolated hunters from the ledger and collect structured candidate results with [HUNTING.md](HUNTING.md), [ATTACK-CLASSES.md](ATTACK-CLASSES.md), and the selected domain companions.
|
||||
3. **Candidate validation** — consolidate fingerprints and give every candidate to a fresh source verifier as defined in [VALIDATION-AND-REPORTING.md](VALIDATION-AND-REPORTING.md).
|
||||
4. **Structured output** — write all final `confirmed`, `needs_validation`, and `rejected` records to `findings.json`; validate it with `report-schema.json` and `validate-findings.cjs`, and validate the coverage claim with `validate-coverage-ledger.cjs`.
|
||||
5. **Independent record verification** — use fresh agents to verify final source claims and reconcile corrections or state changes.
|
||||
6. **Target-neutral report** — derive `REPORT.md`, `FINDINGS-DETAIL.md`, and `NEEDS-VALIDATION.md` from the final records, with no live-probe instructions.
|
||||
|
||||
Do not end the run before one of exactly two terminal states: (a) all Phase 6 artifacts are written and both validators pass, or (b) `run_status: "incomplete"` is recorded with its exact reason and the gap is disclosed in the report. Never stop mid-phase.
|
||||
|
||||
## Anti-patterns
|
||||
|
||||
1. Checklist deviations presented as vulnerabilities.
|
||||
2. Defense-in-depth advice with no reachable boundary violation.
|
||||
3. Live or shared-environment testing where bounded local evidence is insufficient.
|
||||
4. Guessing provider, proxy, browser, identity, or deployment behavior not present in source.
|
||||
5. Treating intended same-principal authority or self-impact as a cross-boundary result.
|
||||
6. Reporting a parser or runtime effect stronger than the observed effect.
|
||||
7. Emitting prose-only hunter results that cannot be deduplicated or verified.
|
||||
8. Re-reporting carried same-source prior confirmed records or using them as exemplars that anchor the hunt.
|
||||
9. Assigning severity to `needs_validation` records.
|
||||
10. Writing the report before independent verification or letting prose and JSON disagree.
|
||||
73
.agents/skills/security-audit/SUPPLY-CHAIN-AND-RELEASE.md
Normal file
73
.agents/skills/security-audit/SUPPLY-CHAIN-AND-RELEASE.md
Normal file
@@ -0,0 +1,73 @@
|
||||
# Supply Chain and Release Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when the target resolves dependencies, builds from untrusted contributions, runs CI, creates release artifacts, signs or promotes builds, loads plugins, or updates deployed software. This domain covers trust handoffs from source and dependency to the artifact a user runs. Use `MEMORY-SAFETY-AND-BINARY.md` for flaws inside a local binary loader and `CLOUD-AND-DEPLOYMENT.md` for runtime workload authority.
|
||||
|
||||
Split large targets into dependency resolution, CI isolation, artifact provenance, release authorization, and updater/plugin trust.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- A mutable or known-vulnerable dependency is not a finding by itself. Show who can influence resolution, which build consumes it, and what execution or release boundary follows.
|
||||
- Follow integrity across every handoff: source identity, resolved inputs, build worker, artifact identity, test result, signature/attestation, promotion, and update consumer.
|
||||
- CI configuration is authorization code. Establish which event triggered a workflow, whose code runs, which secrets and tokens exist, and what it may publish or mutate.
|
||||
- A checksum fetched from the same untrusted location as the artifact does not establish independent integrity. Identify the trusted root and failure behavior.
|
||||
- Use `confirmed` for in-repo control-flow failures with bounded local validation. Use `needs_validation` for branch protection, hosted-runner, registry, signing-service, or production promotion facts that are not observable.
|
||||
```
|
||||
|
||||
## Dependency and build-input attack classes (subagent_type: `general`)
|
||||
|
||||
**Dependency source and namespace confusion**
|
||||
Resolver configuration can select an unintended public/private namespace, fallback registry, mirror, repository, or source URL. Review package names, source priority, lockfile and checksum use, alternate build files, platform-specific resolution, and first-install versus update behavior.
|
||||
|
||||
**Mutable and unbound build inputs**
|
||||
Builds consume branches, tags, unverified submodules, downloaded tools, generated assets, remote includes, floating CI actions, or container tags whose content can change without source review. Require a lower-trust writer and a path into trusted build output; reproducibility by itself does not prove authenticity.
|
||||
|
||||
**Generated-source and codegen provenance gaps**
|
||||
Schemas, vendored archives, generated clients, localization, documentation examples, or binary blobs produce executable or shipped content without the same review and integrity gate as source. Compare local regeneration with committed output and verify who controls input and generator.
|
||||
|
||||
**Build-context inclusion**
|
||||
Secrets, local configuration, repository metadata, test fixtures, or developer artifacts enter a package or image because the build context and ignore rules exceed intended release inputs. Confirm that the resulting artifact exposes a real credential, private data, or privileged configuration.
|
||||
|
||||
## CI and automation attack classes (subagent_type: `general`)
|
||||
|
||||
**Untrusted code in a privileged workflow**
|
||||
A pull request, issue comment, fork, dependency update, or external event runs contributor-controlled code with protected secrets, write tokens, deployment authority, or a trusted runner. Compare trigger type, checkout ref, approval gate, environment protection, and permission narrowing. Do not assume repository-host defaults that are not in source.
|
||||
|
||||
**Workflow command and expression confusion**
|
||||
Attacker-controlled branch names, commit messages, issue fields, artifact names, matrix values, or generated output enter shell commands, template expressions, paths, or privileged workflow inputs without canonical validation.
|
||||
|
||||
**Cache, artifact, and workspace trust mixing**
|
||||
A lower-trust job can populate a cache, artifact, shared workspace, or output that a higher-trust job later restores and executes or releases. Review cache keys and namespaces, artifact producer identity, digest binding, retention, and whether promotion re-resolves by mutable name.
|
||||
|
||||
**Automation identity overreach**
|
||||
CI jobs receive permissions beyond the operation, repository, environment, or duration needed, and untrusted job inputs can select the affected resource. Missing least privilege alone is hardening; require a reachable privileged action.
|
||||
|
||||
## Release and update attack classes (subagent_type: `general`)
|
||||
|
||||
**Build-to-promotion substitution**
|
||||
Tests, review, signature, and publication refer to mutable tags, filenames, channels, or artifact IDs rather than the same immutable digest. Check every copy, repack, architecture merge, and provenance step between build and release.
|
||||
|
||||
**Release authorization and signing-policy gaps**
|
||||
A release or signature is accepted from the wrong workflow, repository, branch, environment, key role, or threshold. Review identity claims inside attestations and verify the consumer validates them, not just a valid signature. Rotation, expiry, and revocation must fail closed where policy requires.
|
||||
|
||||
**Update metadata and rollback confusion**
|
||||
An updater authenticates payload bytes but not version, product, platform, channel, target path, expiry, or rollback state, or it accepts metadata and payload from different authorized transactions. Verify atomic installation and recovery behavior. A signature API call without policy binding is incomplete.
|
||||
|
||||
**Plugin and extension trust expansion**
|
||||
An extension package gains host authority beyond its declared scope, a lower-trust publisher can replace another publisher's identity, or install/update hooks run before authenticity and capability checks. Intended installation of arbitrary same-user plugins is not a privilege boundary.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Walk backward from a released digest or installed update to every source, generated input, credential, worker, cache, test result, and authorization decision.
|
||||
- Compare untrusted and protected workflow events side by side. Mark each persisted channel crossing between them and require an immutable identity plus producer trust.
|
||||
- Review revoked key, failed download, missing attestation, partial platform release, rollback, and registry outage paths. The failure policy is part of release integrity.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Name the lower-trust actor, controllable source/cache/artifact/metadata, consuming trusted job or updater, and resulting unauthorized publication, code inclusion, secret disclosure, or privileged execution.
|
||||
2. Prove artifact identity across the broken handoff. A different mutable name or unbound digest must reach a real consumer.
|
||||
3. Verify built-in package-manager, repository-host, registry, and signing defaults for the pinned version. Unknown hosted controls require `needs_validation`.
|
||||
4. Keep local validation bounded: use a harmless fixture repository, dummy credential marker, local registry/config, and non-production artifact namespace. Do not publish or alter a real release.
|
||||
5. Return `confirmed` only with a complete source-visible handoff and meaningful result. Return `needs_validation` with the precise branch, runner, registry, signing, or deployment fact an owner must observe.
|
||||
186
.agents/skills/security-audit/VALIDATION-AND-REPORTING.md
Normal file
186
.agents/skills/security-audit/VALIDATION-AND-REPORTING.md
Normal file
@@ -0,0 +1,186 @@
|
||||
# Validation, Structured Output, Verification, and Reporting
|
||||
|
||||
### Phase 3: Independently validate every candidate
|
||||
|
||||
After the clean coverage-critic pass or an explicitly recorded early stop, consolidate Phase 2 candidates and carried same-source prior confirmations by stable fingerprint and root cause. Give every unique proposed `confirmed` and `needs_validation` candidate to a fresh `general` verifier that did not hunt it. A carried prior confirmation follows the same current verification path even though hunters exclude that unchanged root cause. A verifier may read hunter or prior artifacts but must re-read every cited current source location and independently run any decisive check it can reproduce safely.
|
||||
|
||||
Assign each verifier a canonical lowercase unique ID and `<output-dir>/agents/<verifier-id>/scratch/` plus parent-owned `artifacts/`. The verifier writes only to `scratch/` and never writes retained artifacts. It receives only the candidate, its linked coverage-unit checks and artifact paths, architecture facts needed to interpret the path, exact relevant companion validation blocks, the promotion procedure block below, the source/local execution boundary, the `confirmed`, `needs_validation`, and `rejected` branches of `report-schema.json` copied verbatim, and prior records with the same fingerprint. It must not receive another verifier's conclusion.
|
||||
|
||||
#### Candidate-verifier prompt
|
||||
|
||||
```text
|
||||
You did not write this candidate. Try to refute it from repository source and bounded
|
||||
local evidence. Do not contact deployed endpoints or external/shared services. Run
|
||||
target-controlled code only inside the approved OS-enforced sandbox: no external
|
||||
network, empty allowlisted environment, read-only target and tools, scratch-only
|
||||
writes, and explicit low resource and wall-clock limits. If any control is unavailable,
|
||||
do not execute; retain the exact missing capability as a needs_validation blocker.
|
||||
Treat every scratch entry as target-controlled after execution. After the sandbox and
|
||||
all its processes terminate, only trusted parent-side code may promote a predeclared
|
||||
scratch-relative file, following the promotion procedure block included verbatim in
|
||||
this prompt. You and target code never write retained artifacts. If promotion is
|
||||
unavailable or fails, do not use that file as evidence.
|
||||
|
||||
1. Verify every trace and evidence file, positive line number, scope, and description.
|
||||
Confirm the first entry is a real lower-trust entrypoint and the last is the
|
||||
claimed sink or boundary effect.
|
||||
2. Reconstruct the strongest source-visible validation, identity, authorization,
|
||||
normalization, lifecycle, framework, and containment controls on the path.
|
||||
Where the architecture summary names a comparable baseline, note whether it
|
||||
shares the pattern — as calibration, never as grounds to dismiss.
|
||||
3. For a proposed confirmed candidate, independently reproduce the minimum observed
|
||||
result when possible. Verify inputs, interface shape, conditions, and affected
|
||||
dummy principal/resource. Do not infer a stronger result or continue after it.
|
||||
4. Verify that likelihood, impact, confidence, and the proposed source fix match only
|
||||
what the evidence establishes.
|
||||
5. For a proposed needs_validation candidate, decide whether the blocker is genuinely
|
||||
outside source/local observation. If source refutes the trace, reject it. If the
|
||||
missing fact remains decisive, keep needs_validation and make the local and
|
||||
owner-observed plans exact and non-destructive.
|
||||
6. Preserve the fingerprint for the same source-derived root cause across every state.
|
||||
|
||||
Return exactly one JSON object and no surrounding prose:
|
||||
{"decision": "confirmed|needs_validation|rejected", "record": { ... }}
|
||||
where record exactly matches the decision's verdict branch of the schema included
|
||||
in this prompt. A corrected record replaces the hunter's wording.
|
||||
```
|
||||
|
||||
Copy this promotion procedure verbatim into every candidate-verifier prompt:
|
||||
|
||||
```text
|
||||
Artifact promotion procedure (trusted parent-side code only):
|
||||
Reference only for you: the parent performs these steps; you never perform them.
|
||||
|
||||
Before execution, the parent opens and retains trusted, non-inheritable directory
|
||||
descriptors for the agent's scratch/ and artifacts/ roots, and records an allowlist
|
||||
of expected scratch-relative artifact files plus explicit per-file and cumulative
|
||||
byte limits. Never pass those descriptors to the agent or sandbox. After the sandbox
|
||||
and all its processes terminate, trusted parent-side code promotes each allowlisted
|
||||
file separately:
|
||||
|
||||
1. Validate the declared relative path: reject absolute, empty, `.`, `..`, or
|
||||
symlinked components.
|
||||
2. Walk each parent component from the retained scratch-root descriptor with
|
||||
no-follow directory-relative operations; never reopen by path.
|
||||
3. Open the leaf no-follow and nonblocking.
|
||||
4. Verify with `fstat` that it is a regular file with link count exactly one and
|
||||
within the recorded per-file and cumulative byte limits.
|
||||
5. Enforce those limits again while reading from that descriptor.
|
||||
6. Copy exactly the verified size, repeat `fstat`, and reject a changed identity,
|
||||
type, link count, or size.
|
||||
7. For the destination, walk every parent component from the retained
|
||||
artifacts-root descriptor with no-follow directory-relative operations; require
|
||||
each existing component to be a real directory, and create any missing directory
|
||||
exclusively before reopening and verifying it no-follow.
|
||||
8. Create the leaf exclusively without following links, verify that the opened
|
||||
destination is a regular file with link count exactly one, and copy from the
|
||||
verified source descriptor without reopening either path.
|
||||
9. Use equivalent race-safe APIs on non-POSIX systems.
|
||||
10. Never recursively copy or glob scratch, extract an archive into artifacts, or
|
||||
open or promote a symlink, FIFO, socket, device, directory, hard-linked file,
|
||||
changing file, or file that exceeds its bound.
|
||||
11. If any check is unavailable, cannot be enforced, or fails, discard the scratch
|
||||
entry; if it is decisive evidence, retain `needs_validation` with the exact
|
||||
promotion blocker.
|
||||
```
|
||||
|
||||
A verifier can promote `needs_validation` to `confirmed` only after independently establishing the complete path and bounded observed result. Demote proposed confirmation to `needs_validation` when a specific deployment or runtime fact remains unknown. Use `rejected` when source, local behavior, a visible control, missing meaningful impact, or an impossible prerequisite refutes the claim. `needs_validation` is never a parking place for a speculative idea.
|
||||
|
||||
The parent checks that each verifier returned the same fingerprint unless it identified a genuinely different root cause. Merge corrections, record the decision in every linked coverage unit, and ensure there is one final record per fingerprint. Discard a malformed or prose-wrapped verifier result without repairing it; re-run that candidate with a fresh verifier when the budget permits, otherwise it remains an unvalidated ledger candidate under the incomplete-run rule.
|
||||
|
||||
When verifier evidence updates a ledger check, set that check's `agent_id` to the verifier's canonical ID and list its nonempty repository-relative `reviewed_paths`. Keep the unit-level `reviewed_paths` equal to the union across checks. Use `method: "source"` with `artifact: null` for source-only review. Use `method: "local"` only with a file successfully promoted by trusted parent-side code below `agents/<check.agent_id>/artifacts/`. The unit retains its original assignment owner, so independently owned hunter and verifier checks can coexist. For a carried prior record's seeded `planned` unit there is no prior owner: the verifier that re-checks it becomes the unit's assignment owner, and its re-check is the unit's first check, moving the unit to `candidate` with the carried fingerprint.
|
||||
|
||||
If a strict total-agent budget cannot cover every candidate, set the run status to incomplete and follow the deterministic budget rule in `SKILL.md`. An unvalidated candidate remains only in the ledger. It does not enter `findings.json` under any verdict.
|
||||
|
||||
### Phase 4: Write and validate `findings.json`
|
||||
|
||||
The parent writes all independently decided records to `<output-dir>/findings.json`, sorted by fingerprint. Include:
|
||||
|
||||
- `confirmed`: source-grounded vulnerabilities with complete local execution evidence, conditions, specific remediation, likelihood/impact/overall severity, and confidence.
|
||||
- `needs_validation`: source-grounded candidates with an exact unresolved blocker and at least one applicable local or owner-observed deployment plan.
|
||||
- `rejected`: source-grounded candidates disproved during validation, retained so future runs do not repeat the unsupported claim without changed evidence.
|
||||
|
||||
Read `report-schema.json` immediately before writing. It uses `additionalProperties: false`; do not carry hunter wrapper fields into a record. Keep these verdict contracts distinct:
|
||||
|
||||
- A `confirmed` record uses `root_cause`, `intended_behavior`, `conditions`, `execution`, `remediation`, `severity`, and `confidence`. It must not use `claimed_root_cause`, `blockers`, `validation_plan`, or `reason`. `execution` is target-neutral and uses the target's native interface: API/HTTP input, CLI call, library call, message, file fixture, browser action, rendered policy, or local harness as applicable. `observed_result` is nonempty and factual.
|
||||
- A `needs_validation` record uses `claimed_root_cause`, `trace`, `evidence`, `blockers`, and at least one nonempty `validation_plan.local` or `validation_plan.deployment` field. Include both only when both contexts can resolve distinct facts. It must not use severity, execution, remediation, reason, or confirmed root cause.
|
||||
- A `rejected` record uses `claimed_root_cause`, `trace`, `evidence`, and `reason`. It must not use severity, execution, remediation, blockers, validation plan, or confirmed root cause.
|
||||
|
||||
Every record has a stable fingerprint, title, description, and repository-relative source paths. A multi-step trace begins with `entrypoint`, ends with `sink`, and uses `propagation` only between them. One-entry traces use `entrypoint` or `sink`. Overall severity cannot exceed demonstrated impact.
|
||||
|
||||
Run:
|
||||
|
||||
```sh
|
||||
node <skill-dir>/validate-findings.cjs <output-dir>/findings.json
|
||||
node <skill-dir>/validate-coverage-ledger.cjs <output-dir>/coverage-ledger.json
|
||||
```
|
||||
|
||||
Fix every structural and semantic error before continuing. The findings validator rejects input beyond 5 MiB, 1,000 top-level findings, or 64 nesting levels, and caps reported error output at 100 messages. Validator success proves format and ledger consistency only.
|
||||
|
||||
### Phase 5: Verify the final records with fresh eyes
|
||||
|
||||
Launch one fresh `research` verifier per final `confirmed` and `needs_validation` record, in parallel. This verifier checks the structured record, not the hunter write-up, and remains inside source/local boundaries.
|
||||
|
||||
In a `quick` run, Phase 3 and Phase 5 merge: the Phase 3 verifier also performs these record checks and returns the final schema-shaped record, so each candidate gets one fresh independent reviewer instead of two. Every other profile keeps the two passes separate. Never skip independent review of a `confirmed` record in any profile.
|
||||
|
||||
For `confirmed`, require it to check:
|
||||
|
||||
1. Every repository-relative trace/evidence path, line, scope, and described operation.
|
||||
2. Real entry interface and exact local input shape.
|
||||
3. Every condition, parser/policy step, source-visible preventing layer, and observed local result.
|
||||
4. Affected principal/resource and demonstrated impact.
|
||||
5. Severity separation: realistic likelihood, demonstrated impact, overall no greater than impact.
|
||||
6. Remediation strategy and any `code_changes`, including whether the fix enforces the invariant without merely moving trust.
|
||||
|
||||
For `needs_validation`, require it to check:
|
||||
|
||||
1. The source path is real and supports only the `claimed_root_cause` stated.
|
||||
2. Every listed blocker is decisive and not already answerable locally.
|
||||
3. The candidate names a boundary and a possible concrete result rather than a generic concern.
|
||||
4. At least one validation-plan field is present and exact. `local` uses a bounded fixture; `deployment` asks an owner to observe a configuration, identity, route, policy, or runtime fact. Do not invent a plan for an inapplicable context, and never send audit traffic to a deployment.
|
||||
5. The fingerprint matches prior/current records for the same root cause.
|
||||
|
||||
Each verifier returns exactly one JSON object: `{"decision":"verified","fingerprint":"..."}` or `{"decision":"replace","reason":"...","record":{...}}`, with no surrounding prose. A replacement record must match its `confirmed`, `needs_validation`, or `rejected` schema branch. Treat a malformed or prose-wrapped Phase 5 result the same way as in Phase 3: discard it without repairing it and re-run with a fresh verifier when the budget permits.
|
||||
|
||||
Do not apply a Phase 5 replacement as final when it promotes a record to a stronger verdict, including any promotion to `confirmed`, or materially changes the root cause, trace, execution input or observed result, demonstrated impact, or severity. Give that complete replacement to a new independent verifier that did not hunt, perform Phase 3 validation, or propose the Phase 5 replacement. The new verifier rechecks the current source and independently reproduces any decisive local result under the execution boundary, then returns `verified` or another replacement. Apply a material replacement only after this fresh verification. If another material replacement results, repeat with a fresh verifier. If budget or independence is unavailable, remove the disputed record from `findings.json`, keep its ledger unit as an unresolved candidate, and set `run_status: "incomplete"` with an exact `incomplete_reason`. Non-material wording or repository-line corrections may be applied directly when they do not change meaning or evidence.
|
||||
|
||||
After every applied replacement, rerun both validators and update linked ledger decisions. If a final verifier identifies a separate root cause, assign a new fingerprint and send it through independent candidate validation before inclusion. Set `run_status: "complete"` only when every ledger candidate has an independent final disposition and every retained record passes Phase 5.
|
||||
|
||||
Do not verify only `confirmed` records. A misleading `needs_validation` handoff wastes owner time and can preserve a false premise.
|
||||
|
||||
### Phase 6: Produce target-neutral reports from final records
|
||||
|
||||
Only after Phase 5 passes for every record retained in `findings.json`, derive prose from the final records, the ledger, and the hunter `hardening` notes retained in ledger bookkeeping. An incomplete run may report independently verified records, but it must identify each unresolved ledger candidate and must not present it as a finding. The prose files never change a verdict, severity, blocker, or demonstrated impact.
|
||||
|
||||
#### `REPORT.md`
|
||||
|
||||
Write:
|
||||
|
||||
1. Run profile, scope, budget (if set) with agents spent versus planned, source ref, sandboxed source-and-local-only execution statement, prior-run use, and explicit deferred and out-of-scope coverage. Name carried same-source confirmations and changed-source revalidations. A `quick`, scoped, budget-limited, or incomplete run states plainly that it is a partial pass. If candidate validation exhausted a strict budget, state that the run is incomplete and list every unvalidated fingerprint and linked unit; do not describe those candidates as findings. If the budget prevented a mandatory critic, state which critic did not run and make no clean-coverage claim.
|
||||
2. One short security posture summary.
|
||||
3. A confirmed-findings table: severity, title, affected boundary, and one-line observed result.
|
||||
4. Each confirmed finding: repository source location, lower-trust principal, target-native bounded reproduction, conditions, actual result, impact, priority rationale, and smallest source fix.
|
||||
5. A separate `NEEDS VALIDATION` table. Give each lead's title, repository trace, exact blocker, bounded local next step, and safe owner-observed deployment check. Do not assign severity or call it a confirmed vulnerability.
|
||||
6. Separate hardening notes and positive source patterns.
|
||||
7. Coverage summary from the ledger: covered, candidate, blocked, and deferred counts, plus important exclusions and the final critic result.
|
||||
|
||||
Do not describe rejected records as findings. Mention their fingerprints only when they explain a prior disagreement or coverage decision.
|
||||
|
||||
#### `FINDINGS-DETAIL.md`
|
||||
|
||||
For each confirmed `medium`, `high`, or `critical` record, copy the complete source path and target-neutral local reproduction:
|
||||
|
||||
- ordered repository-relative trace and evidence;
|
||||
- dummy attacker/principal and affected dummy resource;
|
||||
- native input, invocation, or fixture and exact bounded instructions;
|
||||
- observed output and the security invariant it proves;
|
||||
- conditions and containment;
|
||||
- source-level remediation and regression case.
|
||||
|
||||
#### `NEEDS-VALIDATION.md`
|
||||
|
||||
For every unresolved record, copy the source trace, verified evidence, exact blocker, affected boundary, and each applicable bounded local or owner-observed resolution plan. Keep these as prioritized leads without severity. Do not turn them into live test guidance or assume the missing deployment fact.
|
||||
|
||||
HTTP is one possible native interface, not the default. A library finding may use a function call, a parser a fixture, a CLI a command, a desktop app an IPC or file action, and infrastructure a locally rendered policy. Do not require an endpoint, external account, or live environment that the target does not have.
|
||||
|
||||
Keep the report proportional to the evidence. A clean run may have zero confirmed records. State that result and the remaining coverage/validation limits without inventing LOW findings.
|
||||
105
.agents/skills/security-audit/WEB-PROTOCOL-AND-AUTH.md
Normal file
105
.agents/skills/security-audit/WEB-PROTOCOL-AND-AUTH.md
Normal file
@@ -0,0 +1,105 @@
|
||||
# HTTP-Protocol and Authentication Hunting
|
||||
|
||||
#### When to use this file
|
||||
|
||||
Reach for this file when the target speaks HTTP at a parsing, caching, browser-authentication, or identity boundary: web applications, APIs, reverse proxies, CDNs, gateways, custom HTTP servers, and services implementing sessions, JWT, OAuth/OIDC, SAML, password recovery, MFA, passkeys, API keys, or mTLS. Use this with `ATTACK-CLASSES.md`: access-control review asks whether a principal may perform an operation; this file asks whether the HTTP or identity layer can confuse which principal, request, assurance level, or token the operation belongs to.
|
||||
|
||||
Pick classes from Phase 1. Split a large target into request framing and cache policy, browser authentication, federated identity, strong authentication and recovery, service credentials, and session lifecycle. A single server behind an unobserved managed proxy has little source-confirmable smuggling surface; a proxy or custom parser has much more.
|
||||
|
||||
## Core discipline (include in every agent prompt for this domain)
|
||||
|
||||
```
|
||||
- Framing and cache findings require two interpretations of the same request, response, or key. Name both components and the exact normalized value on each side.
|
||||
- For every credential, find the signature or secret verification and every binding required for its role: issuer, audience, origin, RP, client, session, principal, resource, assurance, expiry, and one-time state.
|
||||
- Host, Forwarded, X-Forwarded-*, Origin, Referer, redirect targets, callback state, and request-derived URLs are trust decisions. Trace each to the affected identity or response.
|
||||
- A missing header, cookie attribute, MFA prompt, or rate limit is not a finding alone. Require an accepted invalid request, cross-principal impact, assurance downgrade, or credential disclosure.
|
||||
- Classify `confirmed` only from complete source evidence and bounded local request/token tests. Use `needs_validation` when proxy, IdP, browser, certificate, secret, or deployed configuration is required but not visible.
|
||||
```
|
||||
|
||||
## HTTP framing and cache attack classes (subagent_type: `general`)
|
||||
|
||||
**Request framing and desynchronization**
|
||||
Front end and back end disagree on request length or header normalization. Review multiple `Content-Length` values, `Transfer-Encoding`, HTTP/2 or HTTP/3 downgrade, header-name normalization, forbidden connection headers, and CR/LF conversion. Confirm which bytes one component assigns to a request and which bytes its peer assigns to the next request.
|
||||
|
||||
**Web cache poisoning through unkeyed input**
|
||||
A request value changes cached content or security-relevant headers but is absent from the cache key. Compare cache key construction with every response variant, including forwarded host/scheme, selected cookies, query normalization, language/device headers, and authorization state.
|
||||
|
||||
**Cache deception and private-response caching**
|
||||
Cache routing treats a private dynamic path as a public static asset, or caches a response whose identity and authorization inputs are missing from policy. Compare edge cacheability with application route parsing, suffix/path-parameter normalization, and response cache directives.
|
||||
|
||||
**Host and forwarded-header trust**
|
||||
Untrusted host/proxy metadata determines absolute URLs, tenant routing, callbacks, reset links, cache keys, or the client address used by authorization. Confirm who can supply the header and whether trusted ingress removes client-provided copies.
|
||||
|
||||
**Response-header injection**
|
||||
Untrusted data reaches `Location`, `Set-Cookie`, CSP, or another response header with unsafe control characters or normalization. Verify framework rejection before reporting and require a security-relevant response change.
|
||||
|
||||
## Browser-session attack classes (subagent_type: `general`)
|
||||
|
||||
**Ordinary CSRF**
|
||||
A browser sends ambient credentials to a state-changing endpoint that accepts a cross-site request without an effective anti-CSRF token, same-site request binding, or strict Origin/Referer validation. Inventory every cookie-authenticated mutation, including form, JSON-like, multipart, method-override, and legacy routes. SameSite is effective only for the cookie and browser contexts actually used; login CSRF and cross-site subresource requests can have different requirements.
|
||||
|
||||
**Session fixation and invalidation**
|
||||
Session identifiers are not rotated on login, account switch, MFA completion, impersonation, or other privilege changes, or remain valid after logout, password change, revocation, and account disable. Check server sessions, refresh tokens, signed cookies, websocket state, cache copies, and fallback endpoints.
|
||||
|
||||
**Cookie scope and transport**
|
||||
A sensitive cookie has an over-broad `Domain` or `Path`, can cross an insecure transport, or conflicts with a sibling cookie that another component selects differently. Bare missing flags remain hardening notes unless a realistic less-trusted origin, network position, or browser path can gain or replace the credential.
|
||||
|
||||
## Federated-identity attack classes (subagent_type: `general`)
|
||||
|
||||
First establish role. Authorization-server controls such as redirect allowlisting and code issuance do not belong to a relying-party client. Verification and binding defects belong to the component consuming the artifact.
|
||||
|
||||
**JWT verification and claim binding**
|
||||
Check signature verification, server-pinned algorithm and key source, then `exp`, `nbf`, `aud`, and `iss`. Review `kid`, `jku`, and `x5u` as untrusted key selectors, duplicate/header normalization, and decode-without-verify paths. A valid token for another service is invalid here even when signed by a trusted issuer.
|
||||
|
||||
**OAuth/OIDC request and callback binding**
|
||||
Validate exact `redirect_uri` ownership where the target is the authorization server; session-bound `state`; PKCE and authorization-code binding where applicable; ID-token issuer/audience/signature/nonce; and selected-IdP binding in multi-provider flows. Compare initial callback, retry, mobile/deep-link, and account-link routes.
|
||||
|
||||
**SAML signed-object and assertion binding**
|
||||
Ensure the element whose signature is validated is the element used as identity. Review unsigned/fallback paths, safe XML parser configuration, canonicalization differences, and freshness/binding fields such as validity windows, audience/recipient, request correlation, and replay state.
|
||||
|
||||
## MFA, passkey, and account-transition attack classes (subagent_type: `general`)
|
||||
|
||||
**MFA enrollment and assurance downgrade**
|
||||
Enrollment, replacement, disablement, recovery-code generation, trusted-device creation, and fallback login require the intended prior assurance. Check that a valid first factor cannot enroll or replace the second factor without policy-required fresh authentication, and that disabled or stale factors stop authorizing sessions.
|
||||
|
||||
**Step-up binding and bypass**
|
||||
A successful challenge upgrades the wrong session, account, tenant, action, or API request, or an alternate route omits the assurance check. Bind the challenge to principal, current session, assurance target, operation or resource when required, expiry, and one-time completion. Compare UI, API, batch, recovery, and resumed-flow paths.
|
||||
|
||||
**WebAuthn and passkey verification**
|
||||
At registration, bind challenge, RP ID, expected origin, credential, user/userHandle, algorithm, and policy-required user verification to the initiating session. At authentication, verify challenge, RP/origin, credential membership, signature, and intended user presence/verification. Check account-discovery and linking flows for userHandle or credential-to-account confusion. Signature-counter handling is meaningful only when the product treats regressions as a clone signal.
|
||||
|
||||
**Account linking and identity collision**
|
||||
Adding an IdP, passkey, email, phone, device, or external account to an existing account must require a current authenticated session, verified ownership of the new identity, policy-required step-up, and callback state bound to the account that initiated linking. Review unlink/relink and invite-acceptance paths for verified-identifier or tenant collisions.
|
||||
|
||||
**Password reset and broader recovery**
|
||||
Recovery tokens, support/admin recovery, backup codes, device migration, and email or phone change often become the weakest authentication path. Verify token randomness, user/action binding, expiry, one-time state, rate/accounting controls, delivery URL trust, and invalidation of prior tokens and sessions. Different responses that only reveal public account existence are not automatically security findings.
|
||||
|
||||
## API-key and mTLS attack classes (subagent_type: `general`)
|
||||
|
||||
**API-key scope and resource binding**
|
||||
A key authenticates to broader tenants, resources, actions, or environments than its server-side record grants, or request parameters override those bindings. Review key lookup, prefix/full-secret verification, type confusion between publishable and secret keys, scope checks, rotation, revocation caches, and bulk endpoints.
|
||||
|
||||
**API-key exposure and unsafe transport**
|
||||
Keys appear in client bundles, URLs, redirects, logs, error paths, build artifacts, or responses accessible to a lower-trust principal. A public identifier called a key is not a secret. Confirm key type and the authority gained by disclosure.
|
||||
|
||||
**mTLS peer and application-identity confusion**
|
||||
A process trusts client-certificate identity headers from any network peer, verifies a chain but maps attacker-influenceable subject text to an account incorrectly, or accepts a certificate for the wrong trust domain, extended usage, audience, or validity policy. Where a trusted proxy terminates mTLS, verify only that proxy can connect, it removes incoming identity headers, and the backend binds the sanitized identity to the request.
|
||||
|
||||
**Certificate lifecycle fallback**
|
||||
Expired, revoked, missing, or renewal-failed certificates cause silent fallback to bearer-only or anonymous operation, or long-lived pooled connections retain authorization after revocation. Missing deployment revocation data makes the result `needs_validation`; an in-repo fail-open branch is source-confirmable.
|
||||
|
||||
## Universal moves (apply across the above)
|
||||
|
||||
- Walk issue → store → transmit → consume → refresh → revoke for every credential and challenge. Compare normal, error, retry, migration, legacy, and account-switch paths.
|
||||
- Enumerate every door to the same identity and every route to the same sensitive operation. The effective policy is the weakest parallel path, not the most polished UI.
|
||||
- Diff parser, proxy, router, cache, and application normalization side by side. For local validation, feed identical bounded request fixtures into each component rather than sending traffic to a live deployment.
|
||||
- For recovery and linking, draw the account before/after graph. Each edge must name the current principal, proof of the new identity, required assurance, callback/session binding, and revocation effect.
|
||||
|
||||
## Validation rules (apply before reporting ANY finding here)
|
||||
|
||||
1. Apply a source-visibility gate. Proxy chains, edge cache keys, IdP policy, certificate trust, browser cookie behavior, secrets, and deployed auth modes may be outside the repository. Record a precise `needs_validation` candidate instead of asserting missing infrastructure behavior.
|
||||
2. For framing and cache findings, name both components and the divergent parse/key. Confirm cross-request, cross-user, or private-response impact with bounded local fixtures.
|
||||
3. For token, MFA, passkey, account-link, recovery, API-key, and mTLS findings, cite the verification line and missing principal/session/resource/origin/audience/action/assurance binding. Prove the server accepts the invalid transition or credential.
|
||||
4. For CSRF, name the ambient credential, state-changing route, accepted cross-site request shape, browser cookie policy, and missing effective check. Read-only actions and routes requiring a non-ambient bearer token do not qualify.
|
||||
5. Verify framework and library defaults. If version or configuration is unknown, use `needs_validation`; do not turn an unverified critical claim into a lower-severity confirmed finding.
|
||||
6. Return `confirmed` only with a complete source trace and observable unauthorized identity, state, or disclosure. For `needs_validation`, name the missing fact and safe local or owner-observed check that resolves it.
|
||||
461
.agents/skills/security-audit/report-schema.json
Normal file
461
.agents/skills/security-audit/report-schema.json
Normal file
@@ -0,0 +1,461 @@
|
||||
{
|
||||
"$comment": "Top-level contract for findings.json. validate-findings.cjs interprets and checks this schema directly.",
|
||||
"type": "array",
|
||||
"items": {
|
||||
"oneOf": [
|
||||
{
|
||||
"type": "object",
|
||||
"description": "A source-grounded vulnerability that was independently demonstrated.",
|
||||
"properties": {
|
||||
"verdict": {
|
||||
"type": "string",
|
||||
"const": "confirmed"
|
||||
},
|
||||
"fingerprint": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"pattern": "^[A-Za-z0-9][A-Za-z0-9._:/@+-]*$",
|
||||
"description": "A stable source-derived identifier that does not change between validation states."
|
||||
},
|
||||
"title": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"root_cause": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"intended_behavior": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"trace": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": {
|
||||
"type": "string",
|
||||
"enum": ["entrypoint", "propagation", "sink"]
|
||||
},
|
||||
"file": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"line": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
},
|
||||
"scope": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["kind", "file", "line", "scope", "description"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"evidence": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"file": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"line": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["file", "line", "description"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"conditions": {
|
||||
"type": "array",
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": {
|
||||
"type": "string",
|
||||
"enum": ["authentication_level", "authorization_role", "user_interaction", "system_configuration", "network_routing", "environmental_dependency", "data_state", "timing_dependency", "third_party_dependency"]
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["kind", "description"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"execution": {
|
||||
"type": "object",
|
||||
"description": "Target-neutral reproduction in the target's native interface.",
|
||||
"properties": {
|
||||
"attacker_perspective": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"payloads": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "string"
|
||||
}
|
||||
},
|
||||
"instructions": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"items": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"observed_result": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["attacker_perspective", "payloads", "instructions", "observed_result"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"remediation": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"strategy": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"code_changes": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"file_name": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"fixed_code": {
|
||||
"type": "string"
|
||||
}
|
||||
},
|
||||
"required": ["file_name", "fixed_code"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"required": ["strategy"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"severity": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"likelihood": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"score": {
|
||||
"type": "string",
|
||||
"enum": ["informational", "low", "medium", "high", "critical"]
|
||||
},
|
||||
"reason": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["score", "reason"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"impact": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"score": {
|
||||
"type": "string",
|
||||
"enum": ["informational", "low", "medium", "high", "critical"]
|
||||
},
|
||||
"reason": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["score", "reason"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"overall_severity": {
|
||||
"type": "string",
|
||||
"enum": ["informational", "low", "medium", "high", "critical"]
|
||||
}
|
||||
},
|
||||
"required": ["likelihood", "impact", "overall_severity"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"confidence": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"score": {
|
||||
"type": "string",
|
||||
"enum": ["low", "medium", "high"]
|
||||
},
|
||||
"reason": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["score", "reason"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"required": ["verdict", "fingerprint", "title", "description", "root_cause", "intended_behavior", "trace", "evidence", "conditions", "execution", "remediation", "severity", "confidence"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"description": "A source-grounded candidate whose decisive validation is blocked.",
|
||||
"properties": {
|
||||
"verdict": {
|
||||
"type": "string",
|
||||
"const": "needs_validation"
|
||||
},
|
||||
"fingerprint": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"pattern": "^[A-Za-z0-9][A-Za-z0-9._:/@+-]*$"
|
||||
},
|
||||
"title": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"claimed_root_cause": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"trace": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": {
|
||||
"type": "string",
|
||||
"enum": ["entrypoint", "propagation", "sink"]
|
||||
},
|
||||
"file": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"line": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
},
|
||||
"scope": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["kind", "file", "line", "scope", "description"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"evidence": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"file": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"line": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["file", "line", "description"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"blockers": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"validation_plan": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"local": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"deployment": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"required": ["verdict", "fingerprint", "title", "description", "claimed_root_cause", "trace", "evidence", "blockers", "validation_plan"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"description": "A source-grounded candidate refuted during validation.",
|
||||
"properties": {
|
||||
"verdict": {
|
||||
"type": "string",
|
||||
"const": "rejected"
|
||||
},
|
||||
"fingerprint": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"pattern": "^[A-Za-z0-9][A-Za-z0-9._:/@+-]*$"
|
||||
},
|
||||
"title": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"claimed_root_cause": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"trace": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": {
|
||||
"type": "string",
|
||||
"enum": ["entrypoint", "propagation", "sink"]
|
||||
},
|
||||
"file": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"line": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
},
|
||||
"scope": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["kind", "file", "line", "scope", "description"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"evidence": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"file": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"line": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
},
|
||||
"description": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["file", "line", "description"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"reason": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"visibleContent": true
|
||||
}
|
||||
},
|
||||
"required": ["verdict", "fingerprint", "title", "description", "claimed_root_cause", "trace", "evidence", "reason"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
872
.agents/skills/security-audit/validate-coverage-ledger.cjs
Executable file
872
.agents/skills/security-audit/validate-coverage-ledger.cjs
Executable file
@@ -0,0 +1,872 @@
|
||||
#!/usr/bin/env node
|
||||
|
||||
/**
|
||||
* Validates coverage-ledger.json and its canonical coverage IDs.
|
||||
* Usage: node validate-coverage-ledger.cjs <path-to-coverage-ledger.json>
|
||||
*/
|
||||
|
||||
const fs = require("node:fs");
|
||||
const path = require("node:path");
|
||||
const { TextDecoder } = require("node:util");
|
||||
|
||||
// Conservative bounds apply before JSON.parse and again to the parsed document.
|
||||
const LIMITS = Object.freeze({
|
||||
inputBytes: 5 * 1024 * 1024,
|
||||
units: 10000,
|
||||
collectionItems: 1000,
|
||||
objectFields: 1000,
|
||||
nestingDepth: 64,
|
||||
preflightValues: 500000,
|
||||
validationErrors: 100,
|
||||
});
|
||||
const MAX_INPUT_BYTES = LIMITS.inputBytes;
|
||||
const MAX_UNITS = LIMITS.units;
|
||||
const MAX_LIST_ITEMS = LIMITS.collectionItems;
|
||||
const MAX_TEXT_LENGTH = 4096;
|
||||
const REQUIRED_FIELDS = [
|
||||
"coverage_id",
|
||||
"canonical_refs",
|
||||
"surface",
|
||||
"boundary",
|
||||
"subsystem",
|
||||
"attack_class",
|
||||
"starting_paths",
|
||||
"ordinary_attack_class_block",
|
||||
"selected_companion_blocks",
|
||||
"excluded_blocks",
|
||||
"prior_status",
|
||||
"attempts",
|
||||
"wave",
|
||||
"status",
|
||||
"agent_id",
|
||||
"reviewed_paths",
|
||||
"local_checks",
|
||||
"result_fingerprints",
|
||||
"unresolved",
|
||||
];
|
||||
const REF_FIELDS = ["surface", "boundary", "subsystem", "attack_class"];
|
||||
const STATUSES = new Set([
|
||||
"planned",
|
||||
"not_applicable",
|
||||
"out_of_scope",
|
||||
"in_progress",
|
||||
"covered",
|
||||
"candidate",
|
||||
"blocked",
|
||||
"deferred",
|
||||
]);
|
||||
const ATTEMPT_STATUSES = new Set(["covered", "candidate", "blocked"]);
|
||||
const ATTEMPT_FIELDS = [
|
||||
"wave",
|
||||
"status",
|
||||
"agent_id",
|
||||
"reviewed_paths",
|
||||
"local_checks",
|
||||
"result_fingerprints",
|
||||
"unresolved",
|
||||
"reassignment_reason",
|
||||
];
|
||||
const PRIOR_STATUSES = new Set([
|
||||
"new",
|
||||
"prior_confirmed_same_source",
|
||||
"prior_confirmed_changed_source",
|
||||
"prior_needs_validation",
|
||||
"prior_deferred",
|
||||
"prior_blocked",
|
||||
"prior_out_of_scope",
|
||||
"prior_covered_same_source",
|
||||
"prior_covered_changed_source",
|
||||
"prior_rejected_claim_changed",
|
||||
"none",
|
||||
]);
|
||||
const FINGERPRINT_PATTERN = /^[A-Za-z0-9][A-Za-z0-9._:/@+-]*$/;
|
||||
const AGENT_ID_PATTERN = /^[a-z0-9][a-z0-9_-]{0,63}$/;
|
||||
const WINDOWS_RESERVED_AGENT_ID = /^(?:con|prn|aux|nul|com[1-9]|lpt[1-9])$/;
|
||||
const VISIBLE_CONTENT = /[^\p{White_Space}\p{Cc}\p{Cf}\p{Default_Ignorable_Code_Point}]/u;
|
||||
const PATH_FORBIDDEN_CHARACTER = /[\p{Cc}\p{Cf}\p{Zl}\p{Zp}\p{Default_Ignorable_Code_Point}]/u;
|
||||
const WINDOWS_RESERVED_COMPONENT = /^(?:con|prn|aux|nul|clock\$|conin\$|conout\$|com[1-9\u00b9\u00b2\u00b3]|lpt[1-9\u00b9\u00b2\u00b3])(?:\.|$)/iu;
|
||||
const UTF8_DECODER = new TextDecoder("utf-8", { fatal: true, ignoreBOM: true });
|
||||
const UNSAFE_DIAGNOSTIC_CHARACTER = /[\p{Cc}\p{Cf}\p{Cs}\p{Zl}\p{Zp}\p{Default_Ignorable_Code_Point}]/gu;
|
||||
const MAX_DIAGNOSTIC_STRING_LENGTH = 256;
|
||||
|
||||
const hasOwn = (value, key) => Object.prototype.hasOwnProperty.call(value, key);
|
||||
|
||||
class JsonStructureError extends Error {}
|
||||
class SafeInputError extends Error {}
|
||||
|
||||
function escapeUnsafeDiagnosticCharacters(value) {
|
||||
return String(value).replace(UNSAFE_DIAGNOSTIC_CHARACTER, (character) => {
|
||||
const codePoint = character.codePointAt(0);
|
||||
return codePoint <= 0xffff
|
||||
? `\\u${codePoint.toString(16).padStart(4, "0")}`
|
||||
: `\\u{${codePoint.toString(16)}}`;
|
||||
});
|
||||
}
|
||||
|
||||
function safeQuote(value) {
|
||||
let serialized;
|
||||
if (typeof value === "string") {
|
||||
const clipped = value.length > MAX_DIAGNOSTIC_STRING_LENGTH
|
||||
? `${value.slice(0, MAX_DIAGNOSTIC_STRING_LENGTH)}...`
|
||||
: value;
|
||||
serialized = JSON.stringify(clipped);
|
||||
} else if (value === null || typeof value === "boolean") {
|
||||
serialized = String(value);
|
||||
} else if (typeof value === "number" && Number.isFinite(value)) {
|
||||
serialized = String(value);
|
||||
} else {
|
||||
serialized = `"<${Array.isArray(value) ? "array" : typeof value}>"`;
|
||||
}
|
||||
return escapeUnsafeDiagnosticCharacters(serialized);
|
||||
}
|
||||
|
||||
function createErrorList() {
|
||||
const errors = [];
|
||||
Object.defineProperty(errors, "push", {
|
||||
value(...messages) {
|
||||
const remaining = LIMITS.validationErrors - this.length;
|
||||
if (remaining > 0) {
|
||||
Array.prototype.push.apply(this, messages.slice(0, remaining).map(escapeUnsafeDiagnosticCharacters));
|
||||
}
|
||||
return this.length;
|
||||
},
|
||||
});
|
||||
return errors;
|
||||
}
|
||||
|
||||
function isObject(value) {
|
||||
return value !== null && typeof value === "object" && !Array.isArray(value);
|
||||
}
|
||||
|
||||
function preflightJsonText(contents) {
|
||||
const containers = [];
|
||||
let rootState = "value";
|
||||
let inString = false;
|
||||
let escaped = false;
|
||||
let totalValues = 0;
|
||||
|
||||
function fail(message) {
|
||||
throw new JsonStructureError(`input ${message}`);
|
||||
}
|
||||
|
||||
function currentContainer() {
|
||||
return containers[containers.length - 1];
|
||||
}
|
||||
|
||||
function countValue() {
|
||||
totalValues++;
|
||||
if (totalValues > LIMITS.preflightValues) {
|
||||
fail(`exceeds ${LIMITS.preflightValues} total value limit`);
|
||||
}
|
||||
}
|
||||
|
||||
function beginValue() {
|
||||
const container = currentContainer();
|
||||
if (!container) {
|
||||
if (rootState !== "value") fail("has malformed JSON structure");
|
||||
rootState = "end";
|
||||
} else if (container.type === "array") {
|
||||
if (container.state !== "firstValueOrEnd" && container.state !== "value") {
|
||||
fail("has malformed JSON array structure");
|
||||
}
|
||||
container.items++;
|
||||
const limit = container.topLevel ? MAX_UNITS : LIMITS.collectionItems;
|
||||
if (container.items > limit) {
|
||||
fail(container.topLevel
|
||||
? `exceeds ${limit} top-level unit limit`
|
||||
: `exceeds ${limit} item array limit`);
|
||||
}
|
||||
container.state = "commaOrEnd";
|
||||
} else {
|
||||
if (container.state !== "value") fail("has malformed JSON object structure");
|
||||
container.state = "commaOrEnd";
|
||||
}
|
||||
countValue();
|
||||
}
|
||||
|
||||
function beginString() {
|
||||
const container = currentContainer();
|
||||
if (container && container.type === "object" &&
|
||||
(container.state === "firstKeyOrEnd" || container.state === "key")) {
|
||||
container.fields++;
|
||||
if (container.fields > LIMITS.objectFields) {
|
||||
fail(`exceeds ${LIMITS.objectFields} field object limit`);
|
||||
}
|
||||
container.state = "colon";
|
||||
} else {
|
||||
beginValue();
|
||||
}
|
||||
inString = true;
|
||||
}
|
||||
|
||||
function beginContainer(type) {
|
||||
beginValue();
|
||||
if (containers.length >= LIMITS.nestingDepth) {
|
||||
fail(`exceeds nesting depth limit ${LIMITS.nestingDepth}`);
|
||||
}
|
||||
containers.push(type === "array"
|
||||
? { type, state: "firstValueOrEnd", items: 0, topLevel: containers.length === 0 }
|
||||
: { type, state: "firstKeyOrEnd", fields: 0 });
|
||||
}
|
||||
|
||||
function closeContainer(type) {
|
||||
const container = currentContainer();
|
||||
if (!container || container.type !== type) fail("has mismatched JSON containers");
|
||||
const canClose = type === "array"
|
||||
? container.state === "firstValueOrEnd" || container.state === "commaOrEnd"
|
||||
: container.state === "firstKeyOrEnd" || container.state === "commaOrEnd";
|
||||
if (!canClose) fail(`has malformed JSON ${type} structure`);
|
||||
containers.pop();
|
||||
}
|
||||
|
||||
function isWhitespace(character) {
|
||||
return character === " " || character === "\t" || character === "\r" || character === "\n";
|
||||
}
|
||||
|
||||
function isTokenDelimiter(character) {
|
||||
return isWhitespace(character) || character === "," || character === ":" ||
|
||||
character === "[" || character === "]" || character === "{" ||
|
||||
character === "}" || character === "\"";
|
||||
}
|
||||
|
||||
for (let index = 0; index < contents.length; index++) {
|
||||
const character = contents[index];
|
||||
if (inString) {
|
||||
if (escaped) {
|
||||
escaped = false;
|
||||
} else if (character === "\\") {
|
||||
escaped = true;
|
||||
} else if (character === "\"") {
|
||||
inString = false;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
if (isWhitespace(character)) continue;
|
||||
|
||||
if (character === "\"") {
|
||||
beginString();
|
||||
} else if (character === "[") {
|
||||
beginContainer("array");
|
||||
} else if (character === "{") {
|
||||
beginContainer("object");
|
||||
} else if (character === "]") {
|
||||
closeContainer("array");
|
||||
} else if (character === "}") {
|
||||
closeContainer("object");
|
||||
} else if (character === ":") {
|
||||
const container = currentContainer();
|
||||
if (!container || container.type !== "object" || container.state !== "colon") {
|
||||
fail("has malformed JSON object structure");
|
||||
}
|
||||
container.state = "value";
|
||||
} else if (character === ",") {
|
||||
const container = currentContainer();
|
||||
if (!container || container.state !== "commaOrEnd") fail("has malformed JSON collection structure");
|
||||
container.state = container.type === "array" ? "value" : "key";
|
||||
} else {
|
||||
beginValue();
|
||||
while (index + 1 < contents.length && !isTokenDelimiter(contents[index + 1])) index++;
|
||||
}
|
||||
}
|
||||
|
||||
if (inString) fail("has an unterminated JSON string");
|
||||
if (containers.length > 0) fail("has truncated JSON structure");
|
||||
if (rootState !== "end") fail("has no JSON value");
|
||||
}
|
||||
|
||||
function preflightDocument(root) {
|
||||
const errors = createErrorList();
|
||||
const stack = [{ value: root, depth: 0, location: "$" }];
|
||||
let visited = 0;
|
||||
|
||||
while (stack.length > 0) {
|
||||
const { value, depth, location } = stack.pop();
|
||||
visited++;
|
||||
if (visited > LIMITS.preflightValues) {
|
||||
errors.push(`$: exceeds ${LIMITS.preflightValues} total values`);
|
||||
return errors;
|
||||
}
|
||||
if (value === null || typeof value !== "object") continue;
|
||||
if (depth >= LIMITS.nestingDepth) {
|
||||
errors.push(`${location}: exceeds nesting depth limit ${LIMITS.nestingDepth}`);
|
||||
return errors;
|
||||
}
|
||||
|
||||
if (Array.isArray(value)) {
|
||||
const limit = location === "$" ? MAX_UNITS : MAX_LIST_ITEMS;
|
||||
if (value.length > limit) {
|
||||
errors.push(`${location}: exceeds ${limit} entries`);
|
||||
return errors;
|
||||
}
|
||||
for (let index = value.length - 1; index >= 0; index--) {
|
||||
stack.push({ value: value[index], depth: depth + 1, location: `${location}[${index}]` });
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
const keys = Object.keys(value);
|
||||
if (keys.length > LIMITS.objectFields) {
|
||||
errors.push(`${location}: exceeds ${LIMITS.objectFields} object fields`);
|
||||
return errors;
|
||||
}
|
||||
for (let index = keys.length - 1; index >= 0; index--) {
|
||||
const key = keys[index];
|
||||
stack.push({ value: value[key], depth: depth + 1, location: `${location}{${index}}` });
|
||||
}
|
||||
}
|
||||
|
||||
return errors;
|
||||
}
|
||||
|
||||
function hasValidUnicodeScalarValues(value) {
|
||||
let index = 0;
|
||||
while (index < value.length) {
|
||||
const first = value.charCodeAt(index++);
|
||||
if (first >= 0xd800 && first <= 0xdbff) {
|
||||
if (index >= value.length) return false;
|
||||
const second = value.charCodeAt(index++);
|
||||
if (second < 0xdc00 || second > 0xdfff) return false;
|
||||
} else if (first >= 0xdc00 && first <= 0xdfff) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
function hasVisibleProse(value) {
|
||||
return hasValidUnicodeScalarValues(value) && VISIBLE_CONTENT.test(value);
|
||||
}
|
||||
|
||||
function isVisibleText(value, maxLength = MAX_TEXT_LENGTH) {
|
||||
return typeof value === "string" &&
|
||||
value.length > 0 &&
|
||||
value.length <= maxLength &&
|
||||
value.trim() === value &&
|
||||
hasVisibleProse(value);
|
||||
}
|
||||
|
||||
function isCanonicalRef(value) {
|
||||
return isVisibleText(value, 1024) &&
|
||||
!PATH_FORBIDDEN_CHARACTER.test(value) &&
|
||||
value.normalize("NFC") === value;
|
||||
}
|
||||
|
||||
function encodeCanonicalRef(value) {
|
||||
if (!isCanonicalRef(value)) throw new TypeError("invalid canonical reference");
|
||||
let encoded = "";
|
||||
for (const byte of Buffer.from(value, "utf8")) {
|
||||
const unreserved =
|
||||
(byte >= 0x41 && byte <= 0x5a) ||
|
||||
(byte >= 0x61 && byte <= 0x7a) ||
|
||||
(byte >= 0x30 && byte <= 0x39) ||
|
||||
byte === 0x2d || byte === 0x2e || byte === 0x5f || byte === 0x7e;
|
||||
encoded += unreserved ? String.fromCharCode(byte) : `%${byte.toString(16).toUpperCase().padStart(2, "0")}`;
|
||||
}
|
||||
return encoded;
|
||||
}
|
||||
|
||||
function canonicalCoverageId(refs) {
|
||||
if (!isObject(refs)) throw new TypeError("canonical_refs must be an object");
|
||||
const fields = hasOwn(refs, "lifecycle") ? [...REF_FIELDS, "lifecycle"] : REF_FIELDS;
|
||||
if (Object.keys(refs).length !== fields.length || fields.some((field) => !hasOwn(refs, field))) {
|
||||
throw new TypeError("canonical_refs has missing or unexpected fields");
|
||||
}
|
||||
return fields.map((field) => encodeCanonicalRef(refs[field])).join("::");
|
||||
}
|
||||
|
||||
function isSafeRelativePath(value) {
|
||||
if (typeof value !== "string" || value.length === 0 || !hasValidUnicodeScalarValues(value) || value.trim() !== value || PATH_FORBIDDEN_CHARACTER.test(value) || value.includes("\\") || value.includes(":")) return false;
|
||||
if (path.posix.isAbsolute(value) || path.win32.isAbsolute(value) || /^[A-Za-z]:/.test(value) || value.startsWith("~")) return false;
|
||||
return value.split("/").every((segment) =>
|
||||
segment !== "" &&
|
||||
segment !== "." &&
|
||||
segment !== ".." &&
|
||||
!/[ .]$/u.test(segment) &&
|
||||
!WINDOWS_RESERVED_COMPONENT.test(segment));
|
||||
}
|
||||
|
||||
function isSafeAgentId(value) {
|
||||
return typeof value === "string" &&
|
||||
AGENT_ID_PATTERN.test(value) &&
|
||||
!WINDOWS_RESERVED_AGENT_ID.test(value);
|
||||
}
|
||||
|
||||
function isOwnedArtifactPath(value, agentId) {
|
||||
if (!isSafeAgentId(agentId) || !isSafeRelativePath(value)) return false;
|
||||
const prefix = `agents/${agentId}/artifacts/`;
|
||||
return value.startsWith(prefix) && value.length > prefix.length;
|
||||
}
|
||||
|
||||
function validateStringArray(value, location, errors, options = {}) {
|
||||
const { allowEmpty = true, fingerprint = false, pathValue = false } = options;
|
||||
if (!Array.isArray(value)) {
|
||||
errors.push(`${location}: expected array`);
|
||||
return;
|
||||
}
|
||||
if (!allowEmpty && value.length === 0) errors.push(`${location}: must not be empty`);
|
||||
if (value.length > MAX_LIST_ITEMS) errors.push(`${location}: exceeds ${MAX_LIST_ITEMS} entries`);
|
||||
const seen = new Set();
|
||||
value.slice(0, MAX_LIST_ITEMS).forEach((entry, index) => {
|
||||
const entryLocation = `${location}[${index}]`;
|
||||
const valid = pathValue ? isSafeRelativePath(entry) : isVisibleText(entry);
|
||||
if (!valid) errors.push(`${entryLocation}: invalid ${pathValue ? "repository-relative path" : "text"}`);
|
||||
if (fingerprint && typeof entry === "string" && !FINGERPRINT_PATTERN.test(entry)) {
|
||||
errors.push(`${entryLocation}: invalid fingerprint`);
|
||||
}
|
||||
if (typeof entry === "string" && seen.has(entry)) errors.push(`${entryLocation}: duplicate entry`);
|
||||
if (typeof entry === "string") seen.add(entry);
|
||||
});
|
||||
}
|
||||
|
||||
function validateChecks(value, location, errors) {
|
||||
if (!Array.isArray(value)) {
|
||||
errors.push(`${location}: expected array`);
|
||||
return;
|
||||
}
|
||||
if (value.length > MAX_LIST_ITEMS) errors.push(`${location}: exceeds ${MAX_LIST_ITEMS} entries`);
|
||||
value.slice(0, MAX_LIST_ITEMS).forEach((check, index) => {
|
||||
const base = `${location}[${index}]`;
|
||||
if (!isObject(check)) {
|
||||
errors.push(`${base}: expected object`);
|
||||
return;
|
||||
}
|
||||
for (const field of ["agent_id", "reviewed_paths", "invariant", "method", "result", "artifact"]) {
|
||||
if (!hasOwn(check, field)) errors.push(`${base}: missing required field ${safeQuote(field)}`);
|
||||
}
|
||||
if (!isSafeAgentId(check.agent_id)) errors.push(`${base}.agent_id: expected a canonical lowercase agent ID`);
|
||||
validateStringArray(check.reviewed_paths, `${base}.reviewed_paths`, errors, { allowEmpty: false, pathValue: true });
|
||||
if (!isVisibleText(check.invariant)) errors.push(`${base}.invariant: invalid text`);
|
||||
if (check.method !== "source" && check.method !== "local") errors.push(`${base}.method: expected "source" or "local"`);
|
||||
if (!isVisibleText(check.result)) errors.push(`${base}.result: invalid text`);
|
||||
if (check.method === "source" && check.artifact !== null) {
|
||||
errors.push(`${base}.artifact: source-only check must use null`);
|
||||
} else if (check.method === "local") {
|
||||
if (!isOwnedArtifactPath(check.artifact, check.agent_id)) {
|
||||
errors.push(`${base}.artifact: local check requires an artifact owned by agent ${safeQuote(isSafeAgentId(check.agent_id) ? check.agent_id : "<agent-id>")}`);
|
||||
}
|
||||
} else if (check.method !== "source" && check.method !== "local" && check.artifact !== null && !isSafeRelativePath(check.artifact)) {
|
||||
errors.push(`${base}.artifact: expected null or a safe output-relative path`);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
function validateReviewedPathOwnership(unit, base, errors) {
|
||||
if (!Array.isArray(unit.reviewed_paths) || !Array.isArray(unit.local_checks)) return;
|
||||
const aggregatePaths = new Set(unit.reviewed_paths.filter((value) => typeof value === "string"));
|
||||
const ownedPaths = new Set();
|
||||
for (const check of unit.local_checks) {
|
||||
if (!isObject(check) || !Array.isArray(check.reviewed_paths)) continue;
|
||||
for (const reviewedPath of check.reviewed_paths) {
|
||||
if (typeof reviewedPath === "string") ownedPaths.add(reviewedPath);
|
||||
}
|
||||
}
|
||||
for (const reviewedPath of aggregatePaths) {
|
||||
if (!ownedPaths.has(reviewedPath)) errors.push(`${base}.reviewed_paths: ${safeQuote(reviewedPath)} has no check owner`);
|
||||
}
|
||||
for (const reviewedPath of ownedPaths) {
|
||||
if (!aggregatePaths.has(reviewedPath)) errors.push(`${base}.local_checks: owned path ${safeQuote(reviewedPath)} is absent from aggregate reviewed_paths`);
|
||||
}
|
||||
}
|
||||
|
||||
function validateExcludedBlocks(value, location, errors) {
|
||||
if (!Array.isArray(value)) {
|
||||
errors.push(`${location}: expected array`);
|
||||
return;
|
||||
}
|
||||
if (value.length > MAX_LIST_ITEMS) errors.push(`${location}: exceeds ${MAX_LIST_ITEMS} entries`);
|
||||
const seen = new Set();
|
||||
value.slice(0, MAX_LIST_ITEMS).forEach((entry, index) => {
|
||||
const base = `${location}[${index}]`;
|
||||
if (!isObject(entry)) {
|
||||
errors.push(`${base}: expected object`);
|
||||
return;
|
||||
}
|
||||
if (!isVisibleText(entry.block)) errors.push(`${base}.block: invalid text`);
|
||||
if (!isVisibleText(entry.reason)) errors.push(`${base}.reason: invalid text`);
|
||||
if (typeof entry.block === "string" && seen.has(entry.block)) errors.push(`${base}.block: duplicate entry`);
|
||||
if (typeof entry.block === "string") seen.add(entry.block);
|
||||
});
|
||||
}
|
||||
|
||||
function semanticKey(unit) {
|
||||
return JSON.stringify([
|
||||
unit.surface,
|
||||
unit.boundary,
|
||||
unit.subsystem,
|
||||
unit.attack_class,
|
||||
hasOwn(unit, "lifecycle") ? unit.lifecycle : null,
|
||||
]);
|
||||
}
|
||||
|
||||
function hasValidSemanticFields(unit) {
|
||||
return ["surface", "boundary", "subsystem", "attack_class"].every((field) => isVisibleText(unit[field])) &&
|
||||
(!hasOwn(unit, "lifecycle") || isVisibleText(unit.lifecycle));
|
||||
}
|
||||
|
||||
function requireEmptyArray(unit, field, base, errors) {
|
||||
if (Array.isArray(unit[field]) && unit[field].length > 0) {
|
||||
errors.push(`${base}.${field}: unit with status ${safeQuote(unit.status)} must keep this array empty`);
|
||||
}
|
||||
}
|
||||
|
||||
function requireNonemptyArray(unit, field, base, errors) {
|
||||
if (!Array.isArray(unit[field]) || unit[field].length === 0) {
|
||||
errors.push(`${base}.${field}: unit with status ${safeQuote(unit.status)} requires entries`);
|
||||
}
|
||||
}
|
||||
|
||||
function validateStateInvariants(unit, base, errors) {
|
||||
const emptyEvidence = () => {
|
||||
requireEmptyArray(unit, "reviewed_paths", base, errors);
|
||||
requireEmptyArray(unit, "local_checks", base, errors);
|
||||
};
|
||||
const requireOwner = () => {
|
||||
if (!isSafeAgentId(unit.agent_id)) errors.push(`${base}.agent_id: unit with status ${safeQuote(unit.status)} requires a canonical lowercase agent ID`);
|
||||
};
|
||||
|
||||
if (unit.status !== "candidate") requireEmptyArray(unit, "result_fingerprints", base, errors);
|
||||
|
||||
switch (unit.status) {
|
||||
case "planned":
|
||||
if (unit.agent_id !== null) errors.push(`${base}.agent_id: planned unit must be unassigned`);
|
||||
emptyEvidence();
|
||||
requireEmptyArray(unit, "unresolved", base, errors);
|
||||
break;
|
||||
case "not_applicable":
|
||||
case "out_of_scope":
|
||||
case "deferred":
|
||||
if (unit.agent_id !== null) errors.push(`${base}.agent_id: unit with status ${safeQuote(unit.status)} must be unassigned`);
|
||||
emptyEvidence();
|
||||
requireNonemptyArray(unit, "unresolved", base, errors);
|
||||
break;
|
||||
case "in_progress":
|
||||
requireOwner();
|
||||
emptyEvidence();
|
||||
requireEmptyArray(unit, "unresolved", base, errors);
|
||||
break;
|
||||
case "blocked":
|
||||
requireOwner();
|
||||
requireNonemptyArray(unit, "reviewed_paths", base, errors);
|
||||
requireNonemptyArray(unit, "local_checks", base, errors);
|
||||
requireNonemptyArray(unit, "unresolved", base, errors);
|
||||
break;
|
||||
case "covered":
|
||||
requireOwner();
|
||||
requireNonemptyArray(unit, "reviewed_paths", base, errors);
|
||||
requireNonemptyArray(unit, "local_checks", base, errors);
|
||||
requireEmptyArray(unit, "unresolved", base, errors);
|
||||
break;
|
||||
case "candidate":
|
||||
requireOwner();
|
||||
requireNonemptyArray(unit, "reviewed_paths", base, errors);
|
||||
requireNonemptyArray(unit, "local_checks", base, errors);
|
||||
requireNonemptyArray(unit, "result_fingerprints", base, errors);
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
function validateAttempts(value, unit, base, errors) {
|
||||
if (!Array.isArray(value)) {
|
||||
errors.push(`${base}.attempts: expected array`);
|
||||
return;
|
||||
}
|
||||
if (value.length > MAX_LIST_ITEMS) errors.push(`${base}.attempts: exceeds ${MAX_LIST_ITEMS} entries`);
|
||||
|
||||
const priorOwners = new Set();
|
||||
const priorArtifacts = new Set();
|
||||
let previousWave = 0;
|
||||
value.slice(0, MAX_LIST_ITEMS).forEach((attempt, index) => {
|
||||
const attemptBase = `${base}.attempts[${index}]`;
|
||||
if (!isObject(attempt)) {
|
||||
errors.push(`${attemptBase}: expected object`);
|
||||
return;
|
||||
}
|
||||
for (const field of ATTEMPT_FIELDS) {
|
||||
if (!hasOwn(attempt, field)) errors.push(`${attemptBase}: missing required field ${safeQuote(field)}`);
|
||||
}
|
||||
if (!Number.isInteger(attempt.wave) || attempt.wave < 1) {
|
||||
errors.push(`${attemptBase}.wave: expected a positive integer`);
|
||||
} else {
|
||||
if (attempt.wave <= previousWave) errors.push(`${attemptBase}.wave: archived attempt waves must be strictly increasing`);
|
||||
if (Number.isInteger(unit.wave) && attempt.wave >= unit.wave) {
|
||||
errors.push(`${attemptBase}.wave: archived attempt wave must precede current wave ${safeQuote(unit.wave)}`);
|
||||
}
|
||||
previousWave = attempt.wave;
|
||||
}
|
||||
if (!ATTEMPT_STATUSES.has(attempt.status)) {
|
||||
errors.push(`${attemptBase}.status: expected "covered", "candidate", or "blocked"`);
|
||||
}
|
||||
let hasFreshOwner = false;
|
||||
if (!isSafeAgentId(attempt.agent_id)) {
|
||||
errors.push(`${attemptBase}.agent_id: archived attempt requires a canonical lowercase agent ID`);
|
||||
} else if (priorOwners.has(attempt.agent_id)) {
|
||||
errors.push(`${attemptBase}.agent_id: assignment owner must be fresh for each attempt`);
|
||||
} else {
|
||||
hasFreshOwner = true;
|
||||
}
|
||||
validateStringArray(attempt.reviewed_paths, `${attemptBase}.reviewed_paths`, errors, { pathValue: true });
|
||||
validateChecks(attempt.local_checks, `${attemptBase}.local_checks`, errors);
|
||||
validateReviewedPathOwnership(attempt, attemptBase, errors);
|
||||
validateStringArray(attempt.result_fingerprints, `${attemptBase}.result_fingerprints`, errors, { fingerprint: true });
|
||||
validateStringArray(attempt.unresolved, `${attemptBase}.unresolved`, errors);
|
||||
if (!isVisibleText(attempt.reassignment_reason)) errors.push(`${attemptBase}.reassignment_reason: invalid text`);
|
||||
validateStateInvariants(attempt, attemptBase, errors);
|
||||
|
||||
if (Array.isArray(attempt.local_checks)) {
|
||||
attempt.local_checks.forEach((check, checkIndex) => {
|
||||
if (!isObject(check)) return;
|
||||
if (priorOwners.has(check.agent_id)) {
|
||||
errors.push(`${attemptBase}.local_checks[${checkIndex}].agent_id: prior assignment owner evidence must remain in its earlier attempt`);
|
||||
}
|
||||
if (typeof check.artifact === "string" && priorArtifacts.has(check.artifact)) {
|
||||
errors.push(`${attemptBase}.local_checks[${checkIndex}].artifact: artifact from an earlier attempt cannot be reused`);
|
||||
}
|
||||
if (check.method === "local" && typeof check.artifact === "string") priorArtifacts.add(check.artifact);
|
||||
});
|
||||
}
|
||||
if (hasFreshOwner) priorOwners.add(attempt.agent_id);
|
||||
});
|
||||
|
||||
if (isSafeAgentId(unit.agent_id) && priorOwners.has(unit.agent_id)) {
|
||||
errors.push(`${base}.agent_id: current assignment owner must be fresh after reassignment`);
|
||||
}
|
||||
if (Array.isArray(unit.local_checks)) {
|
||||
unit.local_checks.forEach((check, index) => {
|
||||
if (!isObject(check)) return;
|
||||
if (priorOwners.has(check.agent_id)) {
|
||||
errors.push(`${base}.local_checks[${index}].agent_id: prior assignment owner evidence must remain in its archived attempt`);
|
||||
}
|
||||
if (typeof check.artifact === "string" && priorArtifacts.has(check.artifact)) {
|
||||
errors.push(`${base}.local_checks[${index}].artifact: artifact from an archived attempt cannot be reused`);
|
||||
}
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
function collectUnitErrors(unit, index) {
|
||||
const errors = createErrorList();
|
||||
const base = `$[${index}]`;
|
||||
if (!isObject(unit)) return [`${base}: expected object`];
|
||||
|
||||
for (const field of REQUIRED_FIELDS) {
|
||||
if (!hasOwn(unit, field)) errors.push(`${base}: missing required field ${safeQuote(field)}`);
|
||||
}
|
||||
for (const field of ["surface", "boundary", "subsystem", "attack_class"]) {
|
||||
if (!isVisibleText(unit[field])) errors.push(`${base}.${field}: invalid text`);
|
||||
}
|
||||
if (hasOwn(unit, "lifecycle") && !isVisibleText(unit.lifecycle)) errors.push(`${base}.lifecycle: invalid text`);
|
||||
|
||||
let expectedId = null;
|
||||
if (!isObject(unit.canonical_refs)) {
|
||||
errors.push(`${base}.canonical_refs: expected object`);
|
||||
} else {
|
||||
const expectedFields = hasOwn(unit.canonical_refs, "lifecycle") ? [...REF_FIELDS, "lifecycle"] : REF_FIELDS;
|
||||
for (const field of expectedFields) {
|
||||
if (!hasOwn(unit.canonical_refs, field)) {
|
||||
errors.push(`${base}.canonical_refs: missing required field ${safeQuote(field)}`);
|
||||
} else if (!isCanonicalRef(unit.canonical_refs[field])) {
|
||||
errors.push(`${base}.canonical_refs.${field}: invalid canonical reference`);
|
||||
}
|
||||
}
|
||||
if (Object.keys(unit.canonical_refs).some((field) => !expectedFields.includes(field))) {
|
||||
errors.push(`${base}.canonical_refs: contains unexpected fields`);
|
||||
}
|
||||
if (hasOwn(unit, "lifecycle") !== hasOwn(unit.canonical_refs, "lifecycle")) {
|
||||
errors.push(`${base}: lifecycle and canonical_refs.lifecycle must appear together`);
|
||||
}
|
||||
try {
|
||||
expectedId = canonicalCoverageId(unit.canonical_refs);
|
||||
} catch {
|
||||
// The specific reference errors above are more useful.
|
||||
}
|
||||
}
|
||||
if (!isVisibleText(unit.coverage_id, 65536)) {
|
||||
errors.push(`${base}.coverage_id: invalid text`);
|
||||
} else if (expectedId !== null && unit.coverage_id !== expectedId) {
|
||||
errors.push(`${base}.coverage_id: expected canonical ID ${safeQuote(expectedId)}`);
|
||||
}
|
||||
|
||||
validateStringArray(unit.starting_paths, `${base}.starting_paths`, errors, { allowEmpty: false, pathValue: true });
|
||||
if (unit.ordinary_attack_class_block !== null && !isVisibleText(unit.ordinary_attack_class_block)) {
|
||||
errors.push(`${base}.ordinary_attack_class_block: expected null or non-empty text`);
|
||||
}
|
||||
validateStringArray(unit.selected_companion_blocks, `${base}.selected_companion_blocks`, errors);
|
||||
validateExcludedBlocks(unit.excluded_blocks, `${base}.excluded_blocks`, errors);
|
||||
if (Array.isArray(unit.selected_companion_blocks) && Array.isArray(unit.excluded_blocks)) {
|
||||
const selected = new Set(unit.selected_companion_blocks);
|
||||
unit.excluded_blocks.forEach((entry, blockIndex) => {
|
||||
if (isObject(entry) && selected.has(entry.block)) {
|
||||
errors.push(`${base}.excluded_blocks[${blockIndex}].block: block is also selected`);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
if (!PRIOR_STATUSES.has(unit.prior_status)) errors.push(`${base}.prior_status: invalid value ${safeQuote(unit.prior_status)}`);
|
||||
if (!STATUSES.has(unit.status)) errors.push(`${base}.status: invalid value ${safeQuote(unit.status)}`);
|
||||
if (!Number.isInteger(unit.wave) || unit.wave < 1) errors.push(`${base}.wave: expected a positive integer`);
|
||||
if (unit.agent_id !== null && !isSafeAgentId(unit.agent_id)) errors.push(`${base}.agent_id: expected null or a safe agent ID`);
|
||||
|
||||
validateAttempts(unit.attempts, unit, base, errors);
|
||||
validateStringArray(unit.reviewed_paths, `${base}.reviewed_paths`, errors, { pathValue: true });
|
||||
validateChecks(unit.local_checks, `${base}.local_checks`, errors);
|
||||
validateReviewedPathOwnership(unit, base, errors);
|
||||
validateStringArray(unit.result_fingerprints, `${base}.result_fingerprints`, errors, { fingerprint: true });
|
||||
validateStringArray(unit.unresolved, `${base}.unresolved`, errors);
|
||||
|
||||
validateStateInvariants(unit, base, errors);
|
||||
|
||||
return errors;
|
||||
}
|
||||
|
||||
function readFileWithinLimit(file) {
|
||||
const noFollow = fs.constants.O_NOFOLLOW;
|
||||
const nonBlock = fs.constants.O_NONBLOCK;
|
||||
if (!Number.isInteger(noFollow) || noFollow === 0 || !Number.isInteger(nonBlock) || nonBlock === 0) {
|
||||
// Node exposes no race-safe fallback on these platforms, so reject all inputs.
|
||||
throw new SafeInputError("OS no-follow and nonblocking input protection is unavailable");
|
||||
}
|
||||
|
||||
let descriptor;
|
||||
try {
|
||||
descriptor = fs.openSync(file, fs.constants.O_RDONLY | noFollow | nonBlock);
|
||||
} catch (error) {
|
||||
if (error && (error.code === "ELOOP" || error.code === "EMLINK")) throw new SafeInputError("input must not be a symlink");
|
||||
throw error;
|
||||
}
|
||||
try {
|
||||
const stat = fs.fstatSync(descriptor);
|
||||
if (!stat.isFile()) throw new SafeInputError("input must be a regular file");
|
||||
if (stat.size > MAX_INPUT_BYTES) throw new SafeInputError(`input exceeds ${MAX_INPUT_BYTES} byte limit`);
|
||||
|
||||
const chunks = [];
|
||||
const buffer = Buffer.allocUnsafe(64 * 1024);
|
||||
let bytesRead = 0;
|
||||
while (true) {
|
||||
const count = fs.readSync(descriptor, buffer, 0, buffer.length, null);
|
||||
if (count === 0) break;
|
||||
bytesRead += count;
|
||||
if (bytesRead > MAX_INPUT_BYTES) throw new SafeInputError(`input exceeds ${MAX_INPUT_BYTES} byte limit`);
|
||||
chunks.push(Buffer.from(buffer.subarray(0, count)));
|
||||
}
|
||||
try {
|
||||
return UTF8_DECODER.decode(Buffer.concat(chunks, bytesRead));
|
||||
} catch {
|
||||
throw new SafeInputError("input is not valid UTF-8");
|
||||
}
|
||||
} finally {
|
||||
fs.closeSync(descriptor);
|
||||
}
|
||||
}
|
||||
|
||||
function validateDocument(ledger) {
|
||||
const errors = createErrorList();
|
||||
if (!Array.isArray(ledger)) {
|
||||
errors.push("$: expected a top-level array");
|
||||
return errors;
|
||||
}
|
||||
if (ledger.length > MAX_UNITS) {
|
||||
errors.push(`$: exceeds ${MAX_UNITS} coverage units`);
|
||||
return errors;
|
||||
}
|
||||
|
||||
errors.push(...preflightDocument(ledger));
|
||||
if (errors.length > 0) return errors;
|
||||
|
||||
const ids = new Map();
|
||||
const semantics = new Map();
|
||||
let previousId = null;
|
||||
for (let index = 0; index < ledger.length && errors.length < LIMITS.validationErrors; index++) {
|
||||
const unit = ledger[index];
|
||||
errors.push(...collectUnitErrors(unit, index));
|
||||
if (errors.length >= LIMITS.validationErrors) break;
|
||||
if (!isObject(unit) || typeof unit.coverage_id !== "string") continue;
|
||||
|
||||
const key = hasValidSemanticFields(unit) ? semanticKey(unit) : null;
|
||||
if (ids.has(unit.coverage_id)) {
|
||||
const previous = ids.get(unit.coverage_id);
|
||||
const qualifier = key !== null && previous.key !== null && previous.key !== key
|
||||
? "canonical identity collision with different semantic fields"
|
||||
: "duplicate coverage ID";
|
||||
errors.push(`$[${index}].coverage_id: ${qualifier} at $[${previous.index}]`);
|
||||
} else {
|
||||
ids.set(unit.coverage_id, { index, key });
|
||||
}
|
||||
if (key !== null && semantics.has(key) && semantics.get(key).id !== unit.coverage_id) {
|
||||
const previous = semantics.get(key);
|
||||
errors.push(`$[${index}].canonical_refs: semantic tuple already uses coverage ID ${safeQuote(previous.id)} at $[${previous.index}]`);
|
||||
} else if (key !== null) {
|
||||
semantics.set(key, { id: unit.coverage_id, index });
|
||||
}
|
||||
if (previousId !== null && previousId > unit.coverage_id) {
|
||||
errors.push(`$[${index}].coverage_id: units must be sorted lexicographically`);
|
||||
}
|
||||
previousId = unit.coverage_id;
|
||||
}
|
||||
return errors;
|
||||
}
|
||||
|
||||
function run(file) {
|
||||
if (!file) {
|
||||
console.error("Usage: node validate-coverage-ledger.cjs <path-to-coverage-ledger.json>");
|
||||
return 1;
|
||||
}
|
||||
|
||||
let contents;
|
||||
try {
|
||||
contents = readFileWithinLimit(file);
|
||||
} catch (error) {
|
||||
const reason = error instanceof SafeInputError ? error.message : "input could not be opened or read safely";
|
||||
console.error(`Failed to read coverage ledger: ${reason}`);
|
||||
return 1;
|
||||
}
|
||||
|
||||
let ledger;
|
||||
try {
|
||||
preflightJsonText(contents);
|
||||
} catch (error) {
|
||||
const reason = error instanceof JsonStructureError ? error.message : "invalid JSON structure";
|
||||
console.error(`Failed to parse coverage ledger: ${reason}`);
|
||||
return 1;
|
||||
}
|
||||
try {
|
||||
ledger = JSON.parse(contents);
|
||||
} catch {
|
||||
console.error("Failed to parse coverage ledger: invalid JSON syntax");
|
||||
return 1;
|
||||
}
|
||||
|
||||
let errors;
|
||||
try {
|
||||
errors = validateDocument(ledger);
|
||||
} catch {
|
||||
console.error("Failed to validate coverage ledger: unexpected validation error");
|
||||
return 1;
|
||||
}
|
||||
for (const message of errors) console.error("ERROR:", message);
|
||||
if (errors.length > 0) {
|
||||
const cap = errors.length === LIMITS.validationErrors ? `; output capped at ${LIMITS.validationErrors}` : "";
|
||||
console.error(`FAIL: ${errors.length} validation error(s)${cap}`);
|
||||
return 1;
|
||||
}
|
||||
console.log(`PASS: ${ledger.length} coverage units valid`);
|
||||
return 0;
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
LIMITS,
|
||||
PATH_FORBIDDEN_CHARACTER,
|
||||
UNSAFE_DIAGNOSTIC_CHARACTER,
|
||||
VISIBLE_CONTENT,
|
||||
WINDOWS_RESERVED_COMPONENT,
|
||||
canonicalCoverageId,
|
||||
encodeCanonicalRef,
|
||||
hasVisibleProse,
|
||||
isSafeAgentId,
|
||||
isSafeRelativePath,
|
||||
preflightJsonText,
|
||||
readFileWithinLimit,
|
||||
safeQuote,
|
||||
validateDocument,
|
||||
};
|
||||
|
||||
if (require.main === module) process.exit(run(process.argv[2]));
|
||||
740
.agents/skills/security-audit/validate-coverage-ledger.test.cjs
Normal file
740
.agents/skills/security-audit/validate-coverage-ledger.test.cjs
Normal file
@@ -0,0 +1,740 @@
|
||||
const assert = require("node:assert/strict");
|
||||
const fs = require("node:fs");
|
||||
const os = require("node:os");
|
||||
const path = require("node:path");
|
||||
const { spawnSync } = require("node:child_process");
|
||||
const test = require("node:test");
|
||||
const {
|
||||
LIMITS,
|
||||
canonicalCoverageId,
|
||||
encodeCanonicalRef,
|
||||
isSafeAgentId,
|
||||
isSafeRelativePath,
|
||||
preflightJsonText,
|
||||
validateDocument,
|
||||
} = require("./validate-coverage-ledger.cjs");
|
||||
|
||||
const validatorPath = path.join(__dirname, "validate-coverage-ledger.cjs");
|
||||
const CLI_TIMEOUT_MS = 5000;
|
||||
const HOSTILE_CLI_TIMEOUT_MS = 15000;
|
||||
const HAS_SAFE_INPUT_OPEN = Number.isInteger(fs.constants.O_NOFOLLOW) &&
|
||||
fs.constants.O_NOFOLLOW !== 0 &&
|
||||
Number.isInteger(fs.constants.O_NONBLOCK) &&
|
||||
fs.constants.O_NONBLOCK !== 0;
|
||||
|
||||
function unit(overrides = {}) {
|
||||
const canonicalRefs = overrides.canonical_refs || {
|
||||
surface: "src/router.ts#POST /users/:id",
|
||||
boundary: "src/authz.ts#requireOwner",
|
||||
subsystem: "packages/api",
|
||||
attack_class: "ATTACK-CLASSES.md#Access control",
|
||||
};
|
||||
const value = {
|
||||
coverage_id: canonicalCoverageId(canonicalRefs),
|
||||
canonical_refs: canonicalRefs,
|
||||
surface: "Update-user route",
|
||||
boundary: "Object ownership",
|
||||
subsystem: "API",
|
||||
attack_class: "Access control",
|
||||
starting_paths: ["src/router.ts", "src/authz.ts"],
|
||||
ordinary_attack_class_block: "ATTACK-CLASSES.md#Access control",
|
||||
selected_companion_blocks: [],
|
||||
excluded_blocks: [{ block: "WEB-PROTOCOL-AND-AUTH.md#Cache behavior", reason: "The route is not cached." }],
|
||||
prior_status: "new",
|
||||
attempts: [],
|
||||
wave: 1,
|
||||
status: "planned",
|
||||
agent_id: null,
|
||||
reviewed_paths: [],
|
||||
local_checks: [],
|
||||
result_fingerprints: [],
|
||||
unresolved: [],
|
||||
};
|
||||
return Object.assign(value, overrides, { canonical_refs: canonicalRefs });
|
||||
}
|
||||
|
||||
function errorsFor(value) {
|
||||
return validateDocument(value);
|
||||
}
|
||||
|
||||
function runCli(contents, options = {}) {
|
||||
const { nodeArgs = [], timeout = CLI_TIMEOUT_MS } = options;
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-coverage-ledger-"));
|
||||
const ledgerPath = path.join(directory, "coverage-ledger.json");
|
||||
try {
|
||||
fs.writeFileSync(ledgerPath, contents);
|
||||
return spawnSync(process.execPath, [...nodeArgs, validatorPath, ledgerPath], {
|
||||
encoding: "utf8",
|
||||
timeout,
|
||||
});
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
}
|
||||
|
||||
function cliOutput(result) {
|
||||
return `${result.stdout}${result.stderr}`;
|
||||
}
|
||||
|
||||
const TERMINAL_CONTROL_PAYLOAD = "\u001b\u0007\u0085\u202e";
|
||||
const TERMINAL_CONTROL_BYTES = [
|
||||
Buffer.from([0x1b]),
|
||||
Buffer.from([0x07]),
|
||||
Buffer.from("\u0085"),
|
||||
Buffer.from("\u202e"),
|
||||
];
|
||||
|
||||
function assertNoInjectedControlBytes(output) {
|
||||
const bytes = Buffer.isBuffer(output) ? output : Buffer.from(output, "utf8");
|
||||
for (const marker of TERMINAL_CONTROL_BYTES) {
|
||||
assert.equal(bytes.indexOf(marker), -1, `found raw control bytes ${marker.toString("hex")}`);
|
||||
}
|
||||
}
|
||||
|
||||
function sourceCheck(agentId = "hunter-1", overrides = {}) {
|
||||
return {
|
||||
agent_id: agentId,
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
invariant: "The route checks object ownership.",
|
||||
method: "source",
|
||||
result: "The owner check applies before the update.",
|
||||
artifact: null,
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
function localCheck(agentId = "hunter-1", overrides = {}) {
|
||||
return sourceCheck(agentId, {
|
||||
method: "local",
|
||||
result: "The bounded fixture accepted the other owner's object.",
|
||||
artifact: `agents/${agentId}/artifacts/result.txt`,
|
||||
...overrides,
|
||||
});
|
||||
}
|
||||
|
||||
function archivedAttempt(overrides = {}) {
|
||||
const value = {
|
||||
wave: 1,
|
||||
status: "blocked",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
result_fingerprints: [],
|
||||
unresolved: ["The deployed policy is unavailable."],
|
||||
reassignment_reason: "The critic found an unchecked parallel path.",
|
||||
};
|
||||
return Object.assign(value, overrides);
|
||||
}
|
||||
|
||||
test("accepts an empty ledger and complete units", () => {
|
||||
assert.deepEqual(errorsFor([]), []);
|
||||
assert.deepEqual(errorsFor([unit()]), []);
|
||||
|
||||
const missingAttempts = unit();
|
||||
delete missingAttempts.attempts;
|
||||
assert(errorsFor([missingAttempts]).some((error) => error.includes('missing required field "attempts"')));
|
||||
|
||||
const covered = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
});
|
||||
assert.deepEqual(errorsFor([covered]), []);
|
||||
});
|
||||
|
||||
test("accepts a complete ledger through the CLI", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const result = runCli(JSON.stringify([unit()]));
|
||||
assert.equal(result.status, 0, cliOutput(result));
|
||||
assert.match(result.stdout, /PASS: 1 coverage units valid/);
|
||||
});
|
||||
|
||||
test("text preflight ignores structural characters and escapes inside strings", () => {
|
||||
const value = unit({
|
||||
surface: "Route \\ slash [list] {object}, colon: quoted \"value\"",
|
||||
excluded_blocks: [{
|
||||
block: "COMPANION.md#Literal [brackets] {braces}",
|
||||
reason: "The text contains a backslash \\ before an escaped \"quote\".",
|
||||
}],
|
||||
});
|
||||
const contents = JSON.stringify([value]);
|
||||
assert.doesNotThrow(() => preflightJsonText(contents));
|
||||
assert.deepEqual(JSON.parse(contents), [value]);
|
||||
if (HAS_SAFE_INPUT_OPEN) {
|
||||
const result = runCli(contents);
|
||||
assert.equal(result.status, 0, cliOutput(result));
|
||||
}
|
||||
});
|
||||
|
||||
test("text preflight enforces structural cardinality limits", () => {
|
||||
const depthLimit = LIMITS.nestingDepth;
|
||||
assert.doesNotThrow(() => preflightJsonText(`${"[".repeat(depthLimit)}0${"]".repeat(depthLimit)}`));
|
||||
assert.throws(
|
||||
() => preflightJsonText(`${"[".repeat(depthLimit + 1)}0${"]".repeat(depthLimit + 1)}`),
|
||||
/exceeds nesting depth limit 64/,
|
||||
);
|
||||
|
||||
const tooManyUnits = `[${"null,".repeat(LIMITS.units)}null]`;
|
||||
assert.throws(() => preflightJsonText(tooManyUnits), /exceeds 10000 top-level unit limit/);
|
||||
|
||||
const tooManyItems = `[[${"null,".repeat(LIMITS.collectionItems)}null]]`;
|
||||
assert.throws(() => preflightJsonText(tooManyItems), /exceeds 1000 item array limit/);
|
||||
|
||||
const objectFields = Array.from(
|
||||
{ length: LIMITS.objectFields + 1 },
|
||||
(_, index) => `"field${index}":null`,
|
||||
).join(",");
|
||||
assert.throws(() => preflightJsonText(`[{${objectFields}}]`), /exceeds 1000 field object limit/);
|
||||
|
||||
const fullArray = `[${"null,".repeat(LIMITS.collectionItems - 1)}null]`;
|
||||
const arraysNeeded = Math.floor(LIMITS.preflightValues / (LIMITS.collectionItems + 1)) + 1;
|
||||
const tooManyValues = `[${Array.from({ length: arraysNeeded }, () => fullArray).join(",")}]`;
|
||||
assert.throws(() => preflightJsonText(tooManyValues), /exceeds 500000 total value limit/);
|
||||
});
|
||||
|
||||
test("text preflight rejects malformed structural truncation cleanly", () => {
|
||||
assert.throws(() => preflightJsonText("["), /truncated JSON structure/);
|
||||
assert.throws(() => preflightJsonText("[\"unterminated"), /unterminated JSON string/);
|
||||
assert.throws(() => preflightJsonText("[{\"field\":1]"), /mismatched JSON containers/);
|
||||
});
|
||||
|
||||
test("derives collision-free canonical IDs from exact UTF-8 references", () => {
|
||||
assert.equal(encodeCanonicalRef("route:POST /users"), "route%3APOST%20%2Fusers");
|
||||
assert.notEqual(encodeCanonicalRef("route name"), encodeCanonicalRef("route-name"));
|
||||
assert.equal(
|
||||
canonicalCoverageId({ surface: "a", boundary: "b", subsystem: "c", attack_class: "d", lifecycle: "retry" }),
|
||||
"a::b::c::d::retry",
|
||||
);
|
||||
assert.throws(() => encodeCanonicalRef("e\u0301"), /invalid canonical reference/);
|
||||
assert.throws(() => encodeCanonicalRef("bad\u0000ref"), /invalid canonical reference/);
|
||||
assert.throws(() => encodeCanonicalRef("hidden\u200bref"), /invalid canonical reference/);
|
||||
});
|
||||
|
||||
test("rejects noncanonical, duplicate, and colliding IDs", () => {
|
||||
const wrong = unit({ coverage_id: "display-label-slug" });
|
||||
assert(errorsFor([wrong]).some((error) => error.includes("expected canonical ID")));
|
||||
|
||||
const duplicate = unit();
|
||||
assert(errorsFor([duplicate, unit()]).some((error) => error.includes("duplicate coverage ID")));
|
||||
|
||||
const collision = unit();
|
||||
const differentMeaning = unit({ surface: "Delete-user route" });
|
||||
assert(errorsFor([collision, differentMeaning]).some((error) => error.includes("canonical identity collision")));
|
||||
});
|
||||
|
||||
test("requires canonical references to be own properties", () => {
|
||||
const inherited = Object.create(unit().canonical_refs);
|
||||
const value = unit();
|
||||
value.canonical_refs = inherited;
|
||||
assert(errorsFor([value]).some((error) => error.includes("missing required field")));
|
||||
});
|
||||
|
||||
test("rejects aliases for one semantic tuple", () => {
|
||||
const first = unit();
|
||||
const refs = { ...first.canonical_refs, surface: "src/alias.ts#updateUser" };
|
||||
const alias = unit({ canonical_refs: refs });
|
||||
const ledger = [first, alias].sort((left, right) => left.coverage_id.localeCompare(right.coverage_id));
|
||||
assert(errorsFor(ledger).some((error) => error.includes("semantic tuple already uses coverage ID")));
|
||||
});
|
||||
|
||||
test("requires lexicographic order", () => {
|
||||
const secondRefs = {
|
||||
surface: "zzz",
|
||||
boundary: "src/authz.ts#requireOwner",
|
||||
subsystem: "packages/api",
|
||||
attack_class: "ATTACK-CLASSES.md#Access control",
|
||||
};
|
||||
assert(errorsFor([unit({ canonical_refs: secondRefs }), unit()])
|
||||
.some((error) => error.includes("sorted lexicographically")));
|
||||
});
|
||||
|
||||
test("validates assignment block maps", () => {
|
||||
const overlap = unit({
|
||||
selected_companion_blocks: ["AI-AND-LLM.md#Tool calls"],
|
||||
excluded_blocks: [{ block: "AI-AND-LLM.md#Tool calls", reason: "Claimed irrelevant." }],
|
||||
});
|
||||
assert(errorsFor([overlap]).some((error) => error.includes("also selected")));
|
||||
|
||||
const noReason = unit({ excluded_blocks: [{ block: "AI-AND-LLM.md#Tool calls", reason: "" }] });
|
||||
assert(errorsFor([noReason]).some((error) => error.includes("reason")));
|
||||
});
|
||||
|
||||
test("requires owned artifacts for local checks and null artifacts for source checks", () => {
|
||||
const local = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [localCheck()],
|
||||
});
|
||||
assert.deepEqual(errorsFor([local]), []);
|
||||
|
||||
const independentlyVerified = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts", "src/authz.ts"],
|
||||
local_checks: [sourceCheck(), localCheck("verifier-1", { reviewed_paths: ["src/authz.ts"] })],
|
||||
});
|
||||
assert.deepEqual(errorsFor([independentlyVerified]), []);
|
||||
|
||||
const unownedPath = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts", "src/authz.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
});
|
||||
assert(errorsFor([unownedPath]).some((error) => error.includes("has no check owner")));
|
||||
|
||||
for (const [checkAgentId, artifact] of [
|
||||
[null, "agents/hunter-1/artifacts/result.txt"],
|
||||
["hunter-1", null],
|
||||
["hunter-1", "result.txt"],
|
||||
["hunter-1", "agents/hunter-2/artifacts/result.txt"],
|
||||
["../hunter", "agents/../hunter/artifacts/result.txt"],
|
||||
]) {
|
||||
const value = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [localCheck(checkAgentId, { artifact })],
|
||||
});
|
||||
assert.notEqual(errorsFor([value]).length, 0, `${checkAgentId}: ${artifact}`);
|
||||
}
|
||||
|
||||
const unownedSource = unit({
|
||||
status: "blocked",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
unresolved: ["The boundary behavior is not source-visible."],
|
||||
});
|
||||
assert(errorsFor([unownedSource]).some((error) => error.includes("unit with status \"blocked\" requires a canonical lowercase agent ID")));
|
||||
|
||||
const sourceWithArtifact = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck("hunter-1", { artifact: "agents/hunter-1/artifacts/source.txt" })],
|
||||
});
|
||||
assert(errorsFor([sourceWithArtifact]).some((error) => error.includes("source-only check must use null")));
|
||||
});
|
||||
|
||||
test("requires canonical lowercase filesystem-safe agent IDs", () => {
|
||||
for (const value of ["hunter-1", "verifier_2", "a0"]) assert.equal(isSafeAgentId(value), true, value);
|
||||
for (const value of ["Hunter-1", "hunter.1", "hunter-1.", "hunter ", "con", "prn", "aux", "nul", "com1", "lpt9", "../hunter"]) {
|
||||
assert.equal(isSafeAgentId(value), false, value);
|
||||
}
|
||||
|
||||
const caseAlias = unit({ status: "in_progress", agent_id: "Hunter-1" });
|
||||
assert(errorsFor([caseAlias]).some((error) => error.includes("canonical lowercase agent ID")));
|
||||
});
|
||||
|
||||
test("enforces state evidence", () => {
|
||||
assert(errorsFor([unit({ status: "in_progress" })]).some((error) => error.includes("unit with status \"in_progress\" requires")));
|
||||
assert(errorsFor([unit({ status: "blocked" })]).some((error) => error.includes("unresolved")));
|
||||
assert(errorsFor([unit({ status: "candidate" })]).some((error) => error.includes("reviewed_paths")));
|
||||
|
||||
assert.deepEqual(errorsFor([unit({ status: "in_progress", agent_id: "hunter-1" })]), []);
|
||||
const inProgressEvidence = unit({
|
||||
status: "in_progress",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
});
|
||||
assert(errorsFor([inProgressEvidence]).some((error) => error.includes("must keep this array empty")));
|
||||
|
||||
const assignedPlanned = unit({
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
});
|
||||
assert(errorsFor([assignedPlanned]).some((error) => error.includes("planned unit must be unassigned")));
|
||||
|
||||
const candidate = unit({
|
||||
status: "candidate",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck("hunter-1", { invariant: "Ownership is required.", result: "No check exists." })],
|
||||
result_fingerprints: ["src-router-missing-owner-check"],
|
||||
unresolved: ["validation_budget_exhausted"],
|
||||
});
|
||||
assert.deepEqual(errorsFor([candidate]), []);
|
||||
|
||||
const blocked = unit({
|
||||
status: "blocked",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
unresolved: ["The deployed policy is unavailable."],
|
||||
});
|
||||
assert.deepEqual(errorsFor([blocked]), []);
|
||||
const blockedFingerprint = { ...blocked, result_fingerprints: ["forbidden-fingerprint"] };
|
||||
assert(errorsFor([blockedFingerprint]).some((error) => error.includes("result_fingerprints")));
|
||||
|
||||
for (const status of ["not_applicable", "out_of_scope", "deferred"]) {
|
||||
assert.deepEqual(errorsFor([unit({ status, unresolved: ["Reason recorded."] })]), []);
|
||||
const invalid = unit({
|
||||
status,
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
result_fingerprints: ["forbidden-fingerprint"],
|
||||
unresolved: ["Reason recorded."],
|
||||
});
|
||||
const errors = errorsFor([invalid]);
|
||||
assert(errors.some((error) => error.includes("must be unassigned")), status);
|
||||
assert(errors.some((error) => error.includes("reviewed_paths")), status);
|
||||
assert(errors.some((error) => error.includes("result_fingerprints")), status);
|
||||
}
|
||||
|
||||
const coveredFingerprint = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
result_fingerprints: ["forbidden-fingerprint"],
|
||||
});
|
||||
assert(errorsFor([coveredFingerprint]).some((error) => error.includes("result_fingerprints")));
|
||||
|
||||
const coveredUnresolved = unit({
|
||||
status: "covered",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck()],
|
||||
unresolved: ["Unexpected unresolved claim."],
|
||||
});
|
||||
assert(errorsFor([coveredUnresolved]).some((error) => error.includes("unresolved")));
|
||||
});
|
||||
|
||||
test("archives prior evidence when a critic assigns a fresh owner", () => {
|
||||
const reassigned = unit({
|
||||
attempts: [archivedAttempt()],
|
||||
wave: 2,
|
||||
status: "in_progress",
|
||||
agent_id: "hunter-2",
|
||||
});
|
||||
assert.deepEqual(errorsFor([reassigned]), []);
|
||||
|
||||
const finalClosure = unit({
|
||||
attempts: [archivedAttempt()],
|
||||
wave: 2,
|
||||
status: "covered",
|
||||
agent_id: "hunter-2",
|
||||
reviewed_paths: ["src/authz.ts"],
|
||||
local_checks: [sourceCheck("hunter-2", { reviewed_paths: ["src/authz.ts"] })],
|
||||
});
|
||||
assert.deepEqual(errorsFor([finalClosure]), []);
|
||||
});
|
||||
|
||||
test("preserves candidate provenance when reassignment must be deferred", () => {
|
||||
const candidateAttempt = archivedAttempt({
|
||||
status: "candidate",
|
||||
result_fingerprints: ["src-router-missing-owner-check"],
|
||||
unresolved: ["validation_budget_exhausted"],
|
||||
});
|
||||
const deferred = unit({
|
||||
attempts: [candidateAttempt],
|
||||
wave: 2,
|
||||
status: "deferred",
|
||||
unresolved: ["quick_profile_final_critic"],
|
||||
});
|
||||
assert.deepEqual(errorsFor([deferred]), []);
|
||||
});
|
||||
|
||||
test("rejects reassignment owner reuse and evidence mixing", () => {
|
||||
const reusedOwner = unit({
|
||||
attempts: [archivedAttempt()],
|
||||
wave: 2,
|
||||
status: "in_progress",
|
||||
agent_id: "hunter-1",
|
||||
});
|
||||
assert(errorsFor([reusedOwner]).some((error) => error.includes("current assignment owner must be fresh")));
|
||||
|
||||
const mixedEvidence = unit({
|
||||
attempts: [archivedAttempt({ local_checks: [localCheck()] })],
|
||||
wave: 2,
|
||||
status: "covered",
|
||||
agent_id: "hunter-2",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [localCheck()],
|
||||
});
|
||||
const errors = errorsFor([mixedEvidence]);
|
||||
assert(errors.some((error) => error.includes("prior assignment owner evidence must remain")));
|
||||
assert(errors.some((error) => error.includes("artifact from an archived attempt cannot be reused")));
|
||||
|
||||
const mixedHistory = unit({
|
||||
attempts: [
|
||||
archivedAttempt({ local_checks: [localCheck()] }),
|
||||
archivedAttempt({
|
||||
wave: 2,
|
||||
agent_id: "hunter-2",
|
||||
local_checks: [localCheck()],
|
||||
}),
|
||||
],
|
||||
wave: 3,
|
||||
status: "in_progress",
|
||||
agent_id: "hunter-3",
|
||||
});
|
||||
const historyErrors = errorsFor([mixedHistory]);
|
||||
assert(historyErrors.some((error) => error.includes("prior assignment owner evidence must remain in its earlier attempt")));
|
||||
assert(historyErrors.some((error) => error.includes("artifact from an earlier attempt cannot be reused")));
|
||||
|
||||
const unordered = unit({
|
||||
attempts: [archivedAttempt(), archivedAttempt({
|
||||
wave: 1,
|
||||
agent_id: "hunter-2",
|
||||
local_checks: [sourceCheck("hunter-2")],
|
||||
})],
|
||||
wave: 3,
|
||||
status: "in_progress",
|
||||
agent_id: "hunter-3",
|
||||
});
|
||||
assert(errorsFor([unordered]).some((error) => error.includes("strictly increasing")));
|
||||
});
|
||||
|
||||
test("rejects unsafe paths and malformed fingerprints", () => {
|
||||
for (const value of [
|
||||
"/etc/passwd",
|
||||
"../src/file.js",
|
||||
"src/../file.js",
|
||||
"src/con.txt",
|
||||
"src/PRN",
|
||||
"src/AUX.c",
|
||||
"src/NUL",
|
||||
"src/CLOCK$.txt",
|
||||
"src/conin$.txt",
|
||||
"src/conout$",
|
||||
"src/COM1.log",
|
||||
"src/lpt9",
|
||||
"src/COM\u00b9.log",
|
||||
"src/COM\u00b2.log",
|
||||
"src/COM\u00b3.log",
|
||||
"src/lpt\u00b9",
|
||||
"src/lpt\u00b2",
|
||||
"src/lpt\u00b3",
|
||||
"src/file.js.",
|
||||
"C:/src/file.js",
|
||||
"src/file\n.js",
|
||||
"src/file\u0085.js",
|
||||
"src/file\u2028.js",
|
||||
"src/file\u200b.js",
|
||||
"src/file\u034f.js",
|
||||
"src/file\ufe0f.js",
|
||||
]) {
|
||||
assert.equal(isSafeRelativePath(value), false, value);
|
||||
}
|
||||
assert.equal(isSafeRelativePath("src/handler.js"), true);
|
||||
assert.equal(isSafeRelativePath("src/caf\u00e9/handler.js"), true);
|
||||
assert(errorsFor([unit({ starting_paths: ["../src/router.ts"] })]).some((error) => error.includes("repository-relative path")));
|
||||
assert(errorsFor([unit({
|
||||
status: "candidate",
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: ["src/router.ts"],
|
||||
local_checks: [sourceCheck("hunter-1", { invariant: "Ownership is required.", result: "No check exists." })],
|
||||
result_fingerprints: ["not stable"],
|
||||
})]).some((error) => error.includes("invalid fingerprint")));
|
||||
});
|
||||
|
||||
test("rejects format, default-ignorable, and invalid-scalar prose", () => {
|
||||
for (const invisible of ["\u200b", "\u034f", "\ufe0f", "\ud800"]) {
|
||||
assert(errorsFor([unit({ surface: invisible })]).some((error) => error.includes("surface")), JSON.stringify(invisible));
|
||||
assert(errorsFor([unit({
|
||||
excluded_blocks: [{ block: "ATTACK-CLASSES.md#Access control", reason: invisible }],
|
||||
})]).some((error) => error.includes("reason")), JSON.stringify(invisible));
|
||||
}
|
||||
});
|
||||
|
||||
test("quotes input-derived controls in direct validation errors", () => {
|
||||
const invalidStatus = `invalid-${TERMINAL_CONTROL_PAYLOAD}`;
|
||||
const invalidPath = `src/${TERMINAL_CONTROL_PAYLOAD}.js`;
|
||||
const value = unit({
|
||||
status: invalidStatus,
|
||||
agent_id: "hunter-1",
|
||||
reviewed_paths: [invalidPath],
|
||||
local_checks: [sourceCheck()],
|
||||
result_fingerprints: ["force-state-error"],
|
||||
});
|
||||
const output = errorsFor([value]).join("\n");
|
||||
|
||||
assert.match(output, /\$\[0\]\.status/);
|
||||
assert.match(output, /\$\[0\]\.reviewed_paths/);
|
||||
assert.match(output, /\\u001b/);
|
||||
assert.match(output, /\\u0007/);
|
||||
assert.match(output, /\\u0085/);
|
||||
assert.match(output, /\\u202e/);
|
||||
assertNoInjectedControlBytes(output);
|
||||
});
|
||||
|
||||
test("quotes input-derived controls in CLI validation errors", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const value = unit({
|
||||
status: `invalid-${TERMINAL_CONTROL_PAYLOAD}`,
|
||||
result_fingerprints: ["force-state-error"],
|
||||
});
|
||||
const result = runCli(JSON.stringify([value]));
|
||||
|
||||
assert.equal(result.status, 1, cliOutput(result));
|
||||
assert.match(result.stderr, /\$\[0\]\.status/);
|
||||
assert.match(result.stderr, /\\u001b/);
|
||||
assertNoInjectedControlBytes(result.stderr);
|
||||
});
|
||||
|
||||
test("returns a generic syntax error without parser-supplied controls", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const malformed = Buffer.concat([
|
||||
Buffer.from("["),
|
||||
Buffer.from(TERMINAL_CONTROL_PAYLOAD),
|
||||
Buffer.from("]"),
|
||||
]);
|
||||
const result = runCli(malformed);
|
||||
|
||||
assert.equal(result.status, 1, cliOutput(result));
|
||||
assert.equal(result.stderr, "Failed to parse coverage ledger: invalid JSON syntax\n");
|
||||
assertNoInjectedControlBytes(result.stderr);
|
||||
});
|
||||
|
||||
test("does not reflect controls from a failed CLI input path", () => {
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-coverage-ledger-path-"));
|
||||
const missingPath = path.join(directory, `missing-${TERMINAL_CONTROL_PAYLOAD}.json`);
|
||||
try {
|
||||
const result = spawnSync(process.execPath, [validatorPath, missingPath], {
|
||||
encoding: "utf8",
|
||||
timeout: CLI_TIMEOUT_MS,
|
||||
});
|
||||
assert.equal(result.status, 1, cliOutput(result));
|
||||
assert.match(result.stderr, /Failed to read coverage ledger:/);
|
||||
assertNoInjectedControlBytes(result.stderr);
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("rejects invalid UTF-8 through the CLI", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const encoded = Buffer.from(JSON.stringify([unit()]));
|
||||
const marker = Buffer.from("Update-user route");
|
||||
const markerOffset = encoded.indexOf(marker);
|
||||
assert.notEqual(markerOffset, -1);
|
||||
const malformed = Buffer.concat([
|
||||
encoded.subarray(0, markerOffset),
|
||||
Buffer.from([0x80]),
|
||||
encoded.subarray(markerOffset + marker.length),
|
||||
]);
|
||||
|
||||
const result = runCli(malformed);
|
||||
const output = cliOutput(result);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /input is not valid UTF-8/);
|
||||
assert.doesNotMatch(output, /TypeError|stack|at validate-coverage-ledger/i);
|
||||
});
|
||||
|
||||
test("rejects a FIFO through the CLI without blocking", { skip: process.platform === "win32" || !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-coverage-ledger-fifo-"));
|
||||
const fifoPath = path.join(directory, "coverage-ledger.json");
|
||||
try {
|
||||
const created = spawnSync("mkfifo", [fifoPath], { encoding: "utf8", timeout: CLI_TIMEOUT_MS });
|
||||
assert.equal(created.status, 0, cliOutput(created));
|
||||
|
||||
const result = spawnSync(process.execPath, [validatorPath, fifoPath], {
|
||||
encoding: "utf8",
|
||||
timeout: CLI_TIMEOUT_MS,
|
||||
});
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /input must be a regular file/);
|
||||
assert.doesNotMatch(output, /stack|at validate-coverage-ledger/i);
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("rejects a symlink through the CLI", { skip: process.platform === "win32" || !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-coverage-ledger-symlink-"));
|
||||
const targetPath = path.join(directory, "target.json");
|
||||
const symlinkPath = path.join(directory, "coverage-ledger.json");
|
||||
try {
|
||||
fs.writeFileSync(targetPath, JSON.stringify([unit()]));
|
||||
fs.symlinkSync(targetPath, symlinkPath);
|
||||
const result = spawnSync(process.execPath, [validatorPath, symlinkPath], {
|
||||
encoding: "utf8",
|
||||
timeout: CLI_TIMEOUT_MS,
|
||||
});
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /input must not be a symlink/);
|
||||
assert.doesNotMatch(output, /stack|at validate-coverage-ledger/i);
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("rejects deeply nested input without recursion failure", () => {
|
||||
let nested = 0;
|
||||
for (let depth = 0; depth < 20000; depth++) nested = [nested];
|
||||
assert(errorsFor(nested).some((error) => error.includes("exceeds nesting depth limit 64")));
|
||||
if (!HAS_SAFE_INPUT_OPEN) return;
|
||||
|
||||
const result = runCli(`${"[".repeat(20000)}0${"]".repeat(20000)}`);
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /exceeds nesting depth limit 64/);
|
||||
assert.doesNotMatch(output, /RangeError|Maximum call stack|stack|at validate-coverage-ledger/i);
|
||||
});
|
||||
|
||||
test("rejects multi-megabyte nesting under a constrained Node heap", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const openContainers = "[".repeat(2000000);
|
||||
const cases = [
|
||||
openContainers,
|
||||
`${openContainers}0${"]".repeat(2000000)}`,
|
||||
];
|
||||
for (const contents of cases) {
|
||||
const result = runCli(contents, {
|
||||
nodeArgs: ["--max-old-space-size=64"],
|
||||
timeout: HOSTILE_CLI_TIMEOUT_MS,
|
||||
});
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /exceeds nesting depth limit 64/);
|
||||
assert.doesNotMatch(output, /heap out of memory|allocation failed|RangeError|Maximum call stack|stack|at validate-coverage-ledger/i);
|
||||
}
|
||||
});
|
||||
|
||||
test("caps malformed 10000-unit validation output", () => {
|
||||
assert.equal(errorsFor(Array.from({ length: LIMITS.units }, () => null)).length, LIMITS.validationErrors);
|
||||
if (!HAS_SAFE_INPUT_OPEN) return;
|
||||
|
||||
const result = runCli(JSON.stringify(Array.from({ length: LIMITS.units }, () => null)));
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /output capped at 100/);
|
||||
assert(output.length < 20000, `unexpected output length ${output.length}`);
|
||||
assert.doesNotMatch(output, /RangeError|Maximum call stack|stack|at validate-coverage-ledger/i);
|
||||
});
|
||||
|
||||
test("rejects malformed top-level data and excessive unit counts", () => {
|
||||
assert.deepEqual(errorsFor({ units: [] }), ["$: expected a top-level array"]);
|
||||
const tooMany = Array.from({ length: 10001 }, () => null);
|
||||
const errors = errorsFor(tooMany);
|
||||
assert.deepEqual(errors, ["$: exceeds 10000 coverage units"]);
|
||||
|
||||
const oversizedCollection = unit({ extra: Array.from({ length: LIMITS.collectionItems + 1 }, () => null) });
|
||||
assert(errorsFor([oversizedCollection]).some((error) => error.includes("exceeds 1000 entries")));
|
||||
});
|
||||
|
||||
test("accepts a canonical ID derived from near-maximum multibyte references", () => {
|
||||
const canonicalRefs = {
|
||||
surface: "\u6f22".repeat(1024),
|
||||
boundary: "\u00e9".repeat(1024),
|
||||
subsystem: "packages/api",
|
||||
attack_class: "\u6f22".repeat(1023) + "\u00e9",
|
||||
};
|
||||
const value = unit({ canonical_refs: canonicalRefs });
|
||||
assert(value.coverage_id.length > 16384, `coverage_id length ${value.coverage_id.length}`);
|
||||
assert(value.coverage_id.length <= 65536, `coverage_id length ${value.coverage_id.length}`);
|
||||
assert.deepEqual(errorsFor([value]), []);
|
||||
|
||||
if (HAS_SAFE_INPUT_OPEN) {
|
||||
const result = runCli(JSON.stringify([value]));
|
||||
assert.equal(result.status, 0, cliOutput(result));
|
||||
}
|
||||
});
|
||||
773
.agents/skills/security-audit/validate-findings.cjs
Executable file
773
.agents/skills/security-audit/validate-findings.cjs
Executable file
@@ -0,0 +1,773 @@
|
||||
#!/usr/bin/env node
|
||||
|
||||
/**
|
||||
* Validates findings.json against report-schema.json.
|
||||
* Usage: node validate-findings.cjs <path-to-findings.json>
|
||||
*
|
||||
* This is a dependency-free interpreter for the JSON Schema keywords used by
|
||||
* report-schema.json, plus finding-specific checks that are clearer in code.
|
||||
*/
|
||||
|
||||
const fs = require("fs");
|
||||
const path = require("path");
|
||||
const { TextDecoder } = require("util");
|
||||
|
||||
const hasOwn = (value, key) => Object.prototype.hasOwnProperty.call(value, key);
|
||||
const SUPPORTED_TYPES = new Set(["object", "array", "string", "integer", "number", "boolean", "null"]);
|
||||
const SUPPORTED_KEYWORDS = new Set([
|
||||
"$comment",
|
||||
"additionalProperties",
|
||||
"const",
|
||||
"description",
|
||||
"enum",
|
||||
"items",
|
||||
"minimum",
|
||||
"minItems",
|
||||
"minLength",
|
||||
"oneOf",
|
||||
"pattern",
|
||||
"properties",
|
||||
"required",
|
||||
"type",
|
||||
"uniqueItems",
|
||||
"visibleContent",
|
||||
]);
|
||||
const SEVERITY_RANK = new Map([
|
||||
["informational", 0],
|
||||
["low", 1],
|
||||
["medium", 2],
|
||||
["high", 3],
|
||||
["critical", 4],
|
||||
]);
|
||||
const LIMITS = Object.freeze({
|
||||
inputBytes: 5 * 1024 * 1024,
|
||||
nestingDepth: 64,
|
||||
arrayItems: 1000,
|
||||
canonicalKeyBytes: 1024 * 1024,
|
||||
uniqueSetBytes: 5 * 1024 * 1024,
|
||||
validationErrors: 100,
|
||||
});
|
||||
const VISIBLE_CONTENT = /[^\p{White_Space}\p{Cc}\p{Cf}\p{Default_Ignorable_Code_Point}]/u;
|
||||
const PATH_FORBIDDEN_CHARACTER = /[\p{Cc}\p{Cf}\p{Zl}\p{Zp}\p{Default_Ignorable_Code_Point}]/u;
|
||||
const WINDOWS_RESERVED_COMPONENT = /^(?:con|prn|aux|nul|clock\$|conin\$|conout\$|com[1-9\u00b9\u00b2\u00b3]|lpt[1-9\u00b9\u00b2\u00b3])(?:\.|$)/iu;
|
||||
const UTF8_DECODER = new TextDecoder("utf-8", { fatal: true, ignoreBOM: true });
|
||||
const UNSAFE_DIAGNOSTIC_CHARACTER = /[\p{Cc}\p{Cf}\p{Cs}\p{Zl}\p{Zp}\p{Default_Ignorable_Code_Point}]/gu;
|
||||
const MAX_DIAGNOSTIC_STRING_LENGTH = 256;
|
||||
|
||||
class JsonStructureError extends Error {}
|
||||
class SafeInputError extends Error {}
|
||||
|
||||
function escapeUnsafeDiagnosticCharacters(value) {
|
||||
return String(value).replace(UNSAFE_DIAGNOSTIC_CHARACTER, (character) => {
|
||||
const codePoint = character.codePointAt(0);
|
||||
return codePoint <= 0xffff
|
||||
? `\\u${codePoint.toString(16).padStart(4, "0")}`
|
||||
: `\\u{${codePoint.toString(16)}}`;
|
||||
});
|
||||
}
|
||||
|
||||
function safeQuote(value) {
|
||||
let serialized;
|
||||
if (typeof value === "string") {
|
||||
const clipped = value.length > MAX_DIAGNOSTIC_STRING_LENGTH
|
||||
? `${value.slice(0, MAX_DIAGNOSTIC_STRING_LENGTH)}...`
|
||||
: value;
|
||||
serialized = JSON.stringify(clipped);
|
||||
} else if (value === null || typeof value === "boolean") {
|
||||
serialized = String(value);
|
||||
} else if (typeof value === "number" && Number.isFinite(value)) {
|
||||
serialized = String(value);
|
||||
} else {
|
||||
serialized = `"<${Array.isArray(value) ? "array" : typeof value}>"`;
|
||||
}
|
||||
return escapeUnsafeDiagnosticCharacters(serialized);
|
||||
}
|
||||
|
||||
function propertyPath(base, key) {
|
||||
return /^[A-Za-z_][A-Za-z0-9_]*$/.test(key)
|
||||
? `${base}.${key}`
|
||||
: `${base}[${safeQuote(key)}]`;
|
||||
}
|
||||
|
||||
function createErrorList() {
|
||||
const errors = [];
|
||||
Object.defineProperty(errors, "push", {
|
||||
value(...messages) {
|
||||
const remaining = LIMITS.validationErrors - this.length;
|
||||
if (remaining > 0) {
|
||||
Array.prototype.push.apply(this, messages.slice(0, remaining).map(escapeUnsafeDiagnosticCharacters));
|
||||
}
|
||||
return this.length;
|
||||
},
|
||||
});
|
||||
return errors;
|
||||
}
|
||||
|
||||
function typeOf(value) {
|
||||
if (Array.isArray(value)) return "array";
|
||||
if (value === null) return "null";
|
||||
return typeof value;
|
||||
}
|
||||
|
||||
function deepEqual(left, right) {
|
||||
if (left === right) return true;
|
||||
if (typeOf(left) !== typeOf(right)) return false;
|
||||
if (Array.isArray(left)) {
|
||||
return left.length === right.length && left.every((value, index) => deepEqual(value, right[index]));
|
||||
}
|
||||
if (left !== null && typeof left === "object") {
|
||||
const leftKeys = Object.keys(left);
|
||||
const rightKeys = Object.keys(right);
|
||||
return leftKeys.length === rightKeys.length &&
|
||||
leftKeys.every((key) => hasOwn(right, key) && deepEqual(left[key], right[key]));
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
function codePointLength(value) {
|
||||
let length = 0;
|
||||
let index = 0;
|
||||
while (index < value.length) {
|
||||
const first = value.charCodeAt(index++);
|
||||
if (first >= 0xd800 && first <= 0xdbff && index < value.length) {
|
||||
const second = value.charCodeAt(index);
|
||||
if (second >= 0xdc00 && second <= 0xdfff) index++;
|
||||
}
|
||||
length++;
|
||||
}
|
||||
return length;
|
||||
}
|
||||
|
||||
function hasValidUnicodeScalarValues(value) {
|
||||
let index = 0;
|
||||
while (index < value.length) {
|
||||
const first = value.charCodeAt(index++);
|
||||
if (first >= 0xd800 && first <= 0xdbff) {
|
||||
if (index >= value.length) return false;
|
||||
const second = value.charCodeAt(index++);
|
||||
if (second < 0xdc00 || second > 0xdfff) return false;
|
||||
} else if (first >= 0xdc00 && first <= 0xdfff) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
function hasVisibleProse(value) {
|
||||
return hasValidUnicodeScalarValues(value) && VISIBLE_CONTENT.test(value);
|
||||
}
|
||||
|
||||
function canonicalKey(value) {
|
||||
const chunks = [];
|
||||
let bytes = 0;
|
||||
|
||||
function append(chunk) {
|
||||
bytes += Buffer.byteLength(chunk);
|
||||
if (bytes > LIMITS.canonicalKeyBytes) {
|
||||
throw new Error(`canonical key exceeds ${LIMITS.canonicalKeyBytes} byte limit`);
|
||||
}
|
||||
chunks.push(chunk);
|
||||
}
|
||||
|
||||
function encode(item) {
|
||||
const type = typeOf(item);
|
||||
if (type === "null") {
|
||||
append("null");
|
||||
} else if (type === "string") {
|
||||
append(`string:${JSON.stringify(item)}`);
|
||||
} else if (type === "number") {
|
||||
append(`number:${Object.is(item, -0) ? "0" : String(item)}`);
|
||||
} else if (type === "boolean") {
|
||||
append(`boolean:${item ? "true" : "false"}`);
|
||||
} else if (type === "array") {
|
||||
append("array:[");
|
||||
item.forEach((entry, index) => {
|
||||
if (index > 0) append(",");
|
||||
encode(entry);
|
||||
});
|
||||
append("]");
|
||||
} else if (type === "object") {
|
||||
append("object:{");
|
||||
Object.keys(item).sort().forEach((key, index) => {
|
||||
if (index > 0) append(",");
|
||||
append(JSON.stringify(key));
|
||||
append(":");
|
||||
encode(item[key]);
|
||||
});
|
||||
append("}");
|
||||
} else {
|
||||
append(`${type}:${String(item)}`);
|
||||
}
|
||||
}
|
||||
|
||||
encode(value);
|
||||
return { key: chunks.join(""), bytes };
|
||||
}
|
||||
|
||||
function collectDataLimitErrors(value, location = "$data") {
|
||||
const stack = [{ value, location, depth: value !== null && typeof value === "object" ? 1 : 0 }];
|
||||
const seen = new WeakSet();
|
||||
|
||||
while (stack.length > 0) {
|
||||
const current = stack.pop();
|
||||
if (current.value === null || typeof current.value !== "object") continue;
|
||||
if (current.depth > LIMITS.nestingDepth) {
|
||||
return [escapeUnsafeDiagnosticCharacters(`${current.location}: exceeds ${LIMITS.nestingDepth} level nesting depth limit`)];
|
||||
}
|
||||
if (seen.has(current.value)) {
|
||||
return [escapeUnsafeDiagnosticCharacters(`${current.location}: input must not contain repeated or cyclic object references`)];
|
||||
}
|
||||
seen.add(current.value);
|
||||
|
||||
if (Array.isArray(current.value)) {
|
||||
if (current.value.length > LIMITS.arrayItems) {
|
||||
return [escapeUnsafeDiagnosticCharacters(`${current.location}: exceeds ${LIMITS.arrayItems} item array limit`)];
|
||||
}
|
||||
for (let index = current.value.length - 1; index >= 0; index--) {
|
||||
const child = current.value[index];
|
||||
if (child !== null && typeof child === "object") {
|
||||
stack.push({ value: child, location: `${current.location}[${index}]`, depth: current.depth + 1 });
|
||||
}
|
||||
}
|
||||
} else {
|
||||
const keys = Object.keys(current.value);
|
||||
for (let index = keys.length - 1; index >= 0; index--) {
|
||||
const key = keys[index];
|
||||
const child = current.value[key];
|
||||
if (child !== null && typeof child === "object") {
|
||||
stack.push({ value: child, location: propertyPath(current.location, key), depth: current.depth + 1 });
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return [];
|
||||
}
|
||||
|
||||
function collectSchemaErrors(schema, location = "schema") {
|
||||
const errors = createErrorList();
|
||||
|
||||
function check(node, p) {
|
||||
if (node === null || typeof node !== "object" || Array.isArray(node)) {
|
||||
errors.push(`${p}: schema must be an object`);
|
||||
return;
|
||||
}
|
||||
|
||||
for (const key of Object.keys(node)) {
|
||||
if (!SUPPORTED_KEYWORDS.has(key)) errors.push(`${p}: unsupported schema keyword ${safeQuote(key)}`);
|
||||
}
|
||||
|
||||
if (hasOwn(node, "$comment") && typeof node.$comment !== "string") {
|
||||
errors.push(`${p}.$comment: expected string`);
|
||||
}
|
||||
if (hasOwn(node, "description") && typeof node.description !== "string") {
|
||||
errors.push(`${p}.description: expected string`);
|
||||
}
|
||||
if (hasOwn(node, "type") && (!SUPPORTED_TYPES.has(node.type))) {
|
||||
errors.push(`${p}.type: unsupported type ${safeQuote(node.type)}`);
|
||||
}
|
||||
if (hasOwn(node, "properties")) {
|
||||
if (node.properties === null || typeof node.properties !== "object" || Array.isArray(node.properties)) {
|
||||
errors.push(`${p}.properties: expected object`);
|
||||
} else {
|
||||
for (const key of Object.keys(node.properties)) check(node.properties[key], propertyPath(`${p}.properties`, key));
|
||||
}
|
||||
}
|
||||
if (hasOwn(node, "required")) {
|
||||
if (!Array.isArray(node.required) || node.required.some((key) => typeof key !== "string")) {
|
||||
errors.push(`${p}.required: expected an array of strings`);
|
||||
} else if (new Set(node.required).size !== node.required.length) {
|
||||
errors.push(`${p}.required: entries must be unique`);
|
||||
}
|
||||
}
|
||||
if (hasOwn(node, "additionalProperties") && typeof node.additionalProperties !== "boolean") {
|
||||
errors.push(`${p}.additionalProperties: only boolean values are supported`);
|
||||
}
|
||||
if (hasOwn(node, "enum")) {
|
||||
if (!Array.isArray(node.enum) || node.enum.length === 0) {
|
||||
errors.push(`${p}.enum: expected a non-empty array`);
|
||||
} else {
|
||||
const seen = new Set();
|
||||
for (const value of node.enum) {
|
||||
let key;
|
||||
try {
|
||||
key = canonicalKey(value).key;
|
||||
} catch (error) {
|
||||
errors.push(`${p}.enum: ${error.message}`);
|
||||
break;
|
||||
}
|
||||
if (seen.has(key)) {
|
||||
errors.push(`${p}.enum: entries must be unique`);
|
||||
break;
|
||||
}
|
||||
seen.add(key);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (hasOwn(node, "items")) check(node.items, `${p}.items`);
|
||||
for (const keyword of ["minItems", "minLength"]) {
|
||||
if (hasOwn(node, keyword) && (!Number.isInteger(node[keyword]) || node[keyword] < 0)) {
|
||||
errors.push(`${p}.${keyword}: expected a non-negative integer`);
|
||||
}
|
||||
}
|
||||
if (hasOwn(node, "minimum") && (typeof node.minimum !== "number" || !Number.isFinite(node.minimum))) {
|
||||
errors.push(`${p}.minimum: expected a finite number`);
|
||||
}
|
||||
if (hasOwn(node, "pattern")) {
|
||||
if (typeof node.pattern !== "string") {
|
||||
errors.push(`${p}.pattern: expected string`);
|
||||
} else {
|
||||
try {
|
||||
new RegExp(node.pattern);
|
||||
} catch (error) {
|
||||
errors.push(`${p}.pattern: invalid regular expression`);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (hasOwn(node, "uniqueItems") && typeof node.uniqueItems !== "boolean") {
|
||||
errors.push(`${p}.uniqueItems: expected boolean`);
|
||||
}
|
||||
if (hasOwn(node, "visibleContent")) {
|
||||
if (typeof node.visibleContent !== "boolean") {
|
||||
errors.push(`${p}.visibleContent: expected boolean`);
|
||||
} else if (node.visibleContent === true && node.type !== "string") {
|
||||
errors.push(`${p}.visibleContent: requires type "string"`);
|
||||
}
|
||||
}
|
||||
if (hasOwn(node, "oneOf")) {
|
||||
if (!Array.isArray(node.oneOf) || node.oneOf.length === 0) {
|
||||
errors.push(`${p}.oneOf: expected a non-empty array`);
|
||||
} else {
|
||||
node.oneOf.forEach((branch, index) => check(branch, `${p}.oneOf[${index}]`));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
check(schema, location);
|
||||
return errors;
|
||||
}
|
||||
|
||||
function findDiscriminator(schema) {
|
||||
if (!hasOwn(schema, "properties") || typeof schema.properties !== "object") return null;
|
||||
for (const key of Object.keys(schema.properties)) {
|
||||
const subSchema = schema.properties[key];
|
||||
if (subSchema && typeof subSchema === "object" && hasOwn(subSchema, "const")) {
|
||||
return { key, value: subSchema.const };
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function validate(value, schema, p, errors) {
|
||||
if (errors.length >= LIMITS.validationErrors) return;
|
||||
if (hasOwn(schema, "oneOf")) {
|
||||
const results = schema.oneOf.map((branch) => collectUnchecked(value, branch, p));
|
||||
const passingIndexes = results
|
||||
.map((branchErrors, index) => branchErrors.length === 0 ? index : -1)
|
||||
.filter((index) => index !== -1);
|
||||
|
||||
if (passingIndexes.length !== 1) {
|
||||
errors.push(`${p}: must match exactly one schema in oneOf; matched ${passingIndexes.length}`);
|
||||
if (passingIndexes.length === 0 && value !== null && typeof value === "object" && !Array.isArray(value)) {
|
||||
const matchingDiscriminators = schema.oneOf
|
||||
.map((branch, index) => ({ discriminator: findDiscriminator(branch), index }))
|
||||
.filter(({ discriminator }) => discriminator && hasOwn(value, discriminator.key) && deepEqual(value[discriminator.key], discriminator.value));
|
||||
if (matchingDiscriminators.length === 1) {
|
||||
errors.push(...results[matchingDiscriminators[0].index]);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (hasOwn(schema, "const") && !deepEqual(value, schema.const)) {
|
||||
errors.push(`${p}: must equal ${safeQuote(schema.const)}, got ${safeQuote(value)}`);
|
||||
}
|
||||
if (hasOwn(schema, "enum") && !schema.enum.some((allowed) => deepEqual(value, allowed))) {
|
||||
const allowed = schema.enum.map(safeQuote).join(", ");
|
||||
errors.push(`${p}: invalid value ${safeQuote(value)} (expected one of ${allowed})`);
|
||||
}
|
||||
|
||||
if (hasOwn(schema, "type") && typeOf(value) !== schema.type && !(schema.type === "integer" && typeOf(value) === "number" && Number.isInteger(value))) {
|
||||
errors.push(`${p}: expected ${schema.type}, got ${typeOf(value)}`);
|
||||
return;
|
||||
}
|
||||
|
||||
if (typeOf(value) === "object") {
|
||||
for (const req of hasOwn(schema, "required") ? schema.required : []) {
|
||||
if (!hasOwn(value, req)) errors.push(`${p}: missing required field ${safeQuote(req)}`);
|
||||
}
|
||||
for (const key of Object.keys(value)) {
|
||||
if (hasOwn(schema, "properties") && hasOwn(schema.properties, key)) {
|
||||
validate(value[key], schema.properties[key], propertyPath(p, key), errors);
|
||||
} else if (hasOwn(schema, "additionalProperties") && schema.additionalProperties === false) {
|
||||
errors.push(`${p}: unexpected field ${safeQuote(key)}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (Array.isArray(value)) {
|
||||
if (hasOwn(schema, "minItems") && value.length < schema.minItems) {
|
||||
errors.push(`${p}: must have at least ${schema.minItems} item(s), got ${value.length}`);
|
||||
}
|
||||
if (hasOwn(schema, "uniqueItems") && schema.uniqueItems === true) {
|
||||
const seen = new Set();
|
||||
let setBytes = 0;
|
||||
for (let i = 0; i < value.length; i++) {
|
||||
let canonical;
|
||||
try {
|
||||
canonical = canonicalKey(value[i]);
|
||||
} catch (error) {
|
||||
errors.push(`${p}[${i}]: ${error.message}`);
|
||||
break;
|
||||
}
|
||||
if (seen.has(canonical.key)) {
|
||||
errors.push(`${p}: items must be unique; duplicate at index ${i}`);
|
||||
continue;
|
||||
}
|
||||
setBytes += canonical.bytes;
|
||||
if (setBytes > LIMITS.uniqueSetBytes) {
|
||||
errors.push(`${p}: canonical uniqueness set exceeds ${LIMITS.uniqueSetBytes} byte limit`);
|
||||
break;
|
||||
}
|
||||
seen.add(canonical.key);
|
||||
}
|
||||
}
|
||||
if (hasOwn(schema, "items")) {
|
||||
value.forEach((item, index) => validate(item, schema.items, `${p}[${index}]`, errors));
|
||||
}
|
||||
}
|
||||
|
||||
if (typeof value === "string") {
|
||||
if (hasOwn(schema, "minLength") && codePointLength(value) < schema.minLength) {
|
||||
errors.push(`${p}: must have at least ${schema.minLength} character(s)`);
|
||||
}
|
||||
if (schema.visibleContent === true) {
|
||||
if (!hasValidUnicodeScalarValues(value)) {
|
||||
errors.push(`${p}: must contain only valid Unicode scalar values`);
|
||||
} else if (!VISIBLE_CONTENT.test(value)) {
|
||||
errors.push(`${p}: must contain a visible character`);
|
||||
}
|
||||
}
|
||||
if (hasOwn(schema, "pattern") && !(new RegExp(schema.pattern).test(value))) {
|
||||
errors.push(`${p}: must match pattern ${JSON.stringify(schema.pattern)}`);
|
||||
}
|
||||
}
|
||||
|
||||
if (typeof value === "number" && hasOwn(schema, "minimum") && value < schema.minimum) {
|
||||
errors.push(`${p}: must be at least ${schema.minimum}, got ${value}`);
|
||||
}
|
||||
}
|
||||
|
||||
function collectUnchecked(value, schema, p) {
|
||||
const errors = createErrorList();
|
||||
validate(value, schema, p, errors);
|
||||
return errors;
|
||||
}
|
||||
|
||||
function collect(value, schema, p = "$data") {
|
||||
const limitErrors = collectDataLimitErrors(value, p);
|
||||
if (limitErrors.length > 0) return limitErrors;
|
||||
return collectUnchecked(value, schema, p);
|
||||
}
|
||||
|
||||
function isSafeRelativeSourcePath(value) {
|
||||
if (typeof value !== "string" || value.length === 0 || !hasValidUnicodeScalarValues(value) || value.trim() !== value || PATH_FORBIDDEN_CHARACTER.test(value) || value.includes("\\") || value.includes(":")) return false;
|
||||
if (path.posix.isAbsolute(value) || path.win32.isAbsolute(value) || /^[A-Za-z]:/.test(value) || value.startsWith("~")) return false;
|
||||
const segments = value.split("/");
|
||||
return segments.every((segment) =>
|
||||
segment !== "" &&
|
||||
segment !== "." &&
|
||||
segment !== ".." &&
|
||||
!/[ .]$/u.test(segment) &&
|
||||
!WINDOWS_RESERVED_COMPONENT.test(segment));
|
||||
}
|
||||
|
||||
function collectFindingSemanticErrors(findings) {
|
||||
const errors = createErrorList();
|
||||
if (!Array.isArray(findings)) return errors;
|
||||
|
||||
const fingerprints = new Map();
|
||||
let previousFingerprint = null;
|
||||
findings.forEach((finding, index) => {
|
||||
if (errors.length >= LIMITS.validationErrors) return;
|
||||
if (!finding || typeof finding !== "object" || Array.isArray(finding)) return;
|
||||
const base = `$[${index}]`;
|
||||
|
||||
if (hasOwn(finding, "fingerprint") && typeof finding.fingerprint === "string") {
|
||||
if (fingerprints.has(finding.fingerprint)) {
|
||||
errors.push(`${base}.fingerprint: duplicate of $[${fingerprints.get(finding.fingerprint)}].fingerprint`);
|
||||
} else {
|
||||
fingerprints.set(finding.fingerprint, index);
|
||||
}
|
||||
if (previousFingerprint !== null && previousFingerprint > finding.fingerprint) {
|
||||
errors.push(`${base}.fingerprint: findings must be sorted lexicographically`);
|
||||
}
|
||||
previousFingerprint = finding.fingerprint;
|
||||
}
|
||||
|
||||
for (const field of ["trace", "evidence"]) {
|
||||
if (!hasOwn(finding, field) || !Array.isArray(finding[field])) continue;
|
||||
finding[field].forEach((entry, entryIndex) => {
|
||||
if (!entry || typeof entry !== "object" || Array.isArray(entry)) return;
|
||||
if (hasOwn(entry, "line") && (!Number.isInteger(entry.line) || entry.line < 1)) {
|
||||
errors.push(`${base}.${field}[${entryIndex}].line: must be a positive integer`);
|
||||
}
|
||||
if (hasOwn(entry, "file") && !isSafeRelativeSourcePath(entry.file)) {
|
||||
errors.push(`${base}.${field}[${entryIndex}].file: must be a safe repository-relative source path`);
|
||||
}
|
||||
});
|
||||
}
|
||||
if (finding.remediation && Array.isArray(finding.remediation.code_changes)) {
|
||||
finding.remediation.code_changes.forEach((change, changeIndex) => {
|
||||
if (change && hasOwn(change, "file_name") && !isSafeRelativeSourcePath(change.file_name)) {
|
||||
errors.push(`${base}.remediation.code_changes[${changeIndex}].file_name: must be a safe repository-relative source path`);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
if (Array.isArray(finding.trace) && finding.trace.length === 1) {
|
||||
const kind = finding.trace[0] && finding.trace[0].kind;
|
||||
if (kind !== "entrypoint" && kind !== "sink") {
|
||||
errors.push(`${base}.trace[0].kind: a one-line trace must be "entrypoint" or "sink"`);
|
||||
}
|
||||
} else if (Array.isArray(finding.trace) && finding.trace.length > 1) {
|
||||
const last = finding.trace.length - 1;
|
||||
if (finding.trace[0] && finding.trace[0].kind !== "entrypoint") {
|
||||
errors.push(`${base}.trace[0].kind: must be "entrypoint", got ${safeQuote(finding.trace[0].kind)}`);
|
||||
}
|
||||
if (finding.trace[last] && finding.trace[last].kind !== "sink") {
|
||||
errors.push(`${base}.trace[${last}].kind: must be "sink", got ${safeQuote(finding.trace[last].kind)}`);
|
||||
}
|
||||
for (let traceIndex = 1; traceIndex < last; traceIndex++) {
|
||||
if (finding.trace[traceIndex] && finding.trace[traceIndex].kind !== "propagation") {
|
||||
errors.push(`${base}.trace[${traceIndex}].kind: must be "propagation", got ${safeQuote(finding.trace[traceIndex].kind)}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const verdict = finding.verdict;
|
||||
if (verdict === "confirmed") {
|
||||
for (const forbidden of ["claimed_root_cause", "blockers", "validation_plan", "reason"]) {
|
||||
if (hasOwn(finding, forbidden)) errors.push(`${base}: confirmed finding must not contain ${safeQuote(forbidden)}`);
|
||||
}
|
||||
if (!finding.execution || typeof finding.execution !== "object" || typeof finding.execution.observed_result !== "string" || !hasVisibleProse(finding.execution.observed_result)) {
|
||||
errors.push(`${base}: confirmed finding requires a visible execution observed_result`);
|
||||
}
|
||||
if (!finding.remediation || typeof finding.remediation !== "object" || typeof finding.remediation.strategy !== "string" || !hasVisibleProse(finding.remediation.strategy)) {
|
||||
errors.push(`${base}: confirmed finding requires visible remediation`);
|
||||
}
|
||||
const overall = finding.severity && finding.severity.overall_severity;
|
||||
const impact = finding.severity && finding.severity.impact && finding.severity.impact.score;
|
||||
if (SEVERITY_RANK.has(overall) && SEVERITY_RANK.has(impact) && SEVERITY_RANK.get(overall) > SEVERITY_RANK.get(impact)) {
|
||||
errors.push(`${base}.severity.overall_severity: cannot exceed demonstrated impact ${safeQuote(impact)}`);
|
||||
}
|
||||
} else if (verdict === "needs_validation") {
|
||||
if (hasOwn(finding, "severity")) errors.push(`${base}: needs_validation finding must not contain "severity"`);
|
||||
for (const forbidden of ["execution", "remediation", "reason", "root_cause"]) {
|
||||
if (hasOwn(finding, forbidden)) errors.push(`${base}: needs_validation finding must not contain ${safeQuote(forbidden)}`);
|
||||
}
|
||||
const plan = finding.validation_plan;
|
||||
const hasLocalPlan = plan && typeof plan.local === "string" && hasVisibleProse(plan.local);
|
||||
const hasDeploymentPlan = plan && typeof plan.deployment === "string" && hasVisibleProse(plan.deployment);
|
||||
if (!hasLocalPlan && !hasDeploymentPlan) {
|
||||
errors.push(`${base}.validation_plan: requires at least one visible local or deployment plan`);
|
||||
}
|
||||
} else if (verdict === "rejected") {
|
||||
for (const forbidden of ["severity", "execution", "remediation", "blockers", "validation_plan", "root_cause"]) {
|
||||
if (hasOwn(finding, forbidden)) errors.push(`${base}: rejected finding must not contain ${safeQuote(forbidden)}`);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
return errors;
|
||||
}
|
||||
|
||||
function validateDocument(findings, schema) {
|
||||
const schemaErrors = collectSchemaErrors(schema);
|
||||
if (schemaErrors.length > 0) return schemaErrors;
|
||||
const limitErrors = collectDataLimitErrors(findings, "$");
|
||||
if (limitErrors.length > 0) return limitErrors;
|
||||
const errors = collectUnchecked(findings, schema, "$");
|
||||
if (errors.length < LIMITS.validationErrors) {
|
||||
errors.push(...collectFindingSemanticErrors(findings));
|
||||
}
|
||||
return errors;
|
||||
}
|
||||
|
||||
function loadSchema(schemaPath) {
|
||||
const schema = JSON.parse(fs.readFileSync(schemaPath, "utf8"));
|
||||
const errors = collectSchemaErrors(schema);
|
||||
if (errors.length > 0) throw new Error(`unsupported or invalid report schema:\n${errors.join("\n")}`);
|
||||
return schema;
|
||||
}
|
||||
|
||||
function readFileWithinLimit(file) {
|
||||
const noFollow = fs.constants.O_NOFOLLOW;
|
||||
const nonBlock = fs.constants.O_NONBLOCK;
|
||||
if (!Number.isInteger(noFollow) || noFollow === 0 || !Number.isInteger(nonBlock) || nonBlock === 0) {
|
||||
throw new SafeInputError("OS no-follow and nonblocking input protection is unavailable");
|
||||
}
|
||||
|
||||
let descriptor;
|
||||
try {
|
||||
descriptor = fs.openSync(file, fs.constants.O_RDONLY | noFollow | nonBlock);
|
||||
} catch (error) {
|
||||
if (error && (error.code === "ELOOP" || error.code === "EMLINK")) {
|
||||
throw new SafeInputError("input must not be a symlink");
|
||||
}
|
||||
throw error;
|
||||
}
|
||||
try {
|
||||
const stat = fs.fstatSync(descriptor);
|
||||
if (!stat.isFile()) {
|
||||
throw new SafeInputError("input must be a regular file");
|
||||
}
|
||||
if (stat.size > LIMITS.inputBytes) {
|
||||
throw new SafeInputError(`input exceeds ${LIMITS.inputBytes} byte limit`);
|
||||
}
|
||||
|
||||
const chunks = [];
|
||||
const buffer = Buffer.allocUnsafe(64 * 1024);
|
||||
let bytesRead = 0;
|
||||
while (true) {
|
||||
const count = fs.readSync(descriptor, buffer, 0, buffer.length, null);
|
||||
if (count === 0) break;
|
||||
bytesRead += count;
|
||||
if (bytesRead > LIMITS.inputBytes) {
|
||||
throw new SafeInputError(`input exceeds ${LIMITS.inputBytes} byte limit`);
|
||||
}
|
||||
chunks.push(Buffer.from(buffer.subarray(0, count)));
|
||||
}
|
||||
try {
|
||||
return UTF8_DECODER.decode(Buffer.concat(chunks, bytesRead));
|
||||
} catch {
|
||||
throw new SafeInputError("input is not valid UTF-8");
|
||||
}
|
||||
} finally {
|
||||
fs.closeSync(descriptor);
|
||||
}
|
||||
}
|
||||
|
||||
function enforceJsonTextLimits(contents) {
|
||||
const containers = [];
|
||||
let inString = false;
|
||||
let escaped = false;
|
||||
|
||||
function markArrayItem() {
|
||||
const container = containers[containers.length - 1];
|
||||
if (!container || container.type !== "array" || !container.expectsItem) return;
|
||||
container.expectsItem = false;
|
||||
container.items++;
|
||||
if (container.items > LIMITS.arrayItems) {
|
||||
throw new JsonStructureError(`input exceeds ${LIMITS.arrayItems} item array limit`);
|
||||
}
|
||||
}
|
||||
|
||||
for (let index = 0; index < contents.length; index++) {
|
||||
const character = contents[index];
|
||||
if (inString) {
|
||||
if (escaped) {
|
||||
escaped = false;
|
||||
} else if (character === "\\") {
|
||||
escaped = true;
|
||||
} else if (character === "\"") {
|
||||
inString = false;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
if (character === "\"") {
|
||||
markArrayItem();
|
||||
inString = true;
|
||||
} else if (character === "[" || character === "{") {
|
||||
markArrayItem();
|
||||
if (containers.length >= LIMITS.nestingDepth) {
|
||||
throw new JsonStructureError(`input exceeds ${LIMITS.nestingDepth} level nesting depth limit`);
|
||||
}
|
||||
containers.push({
|
||||
type: character === "[" ? "array" : "object",
|
||||
expectsItem: character === "[",
|
||||
items: 0,
|
||||
});
|
||||
} else if (character === "]" || character === "}") {
|
||||
containers.pop();
|
||||
} else if (character === ",") {
|
||||
const container = containers[containers.length - 1];
|
||||
if (container && container.type === "array") container.expectsItem = true;
|
||||
} else if (!/\s/.test(character)) {
|
||||
markArrayItem();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function run(file) {
|
||||
if (!file) {
|
||||
console.error("Usage: node validate-findings.cjs <path-to-findings.json>");
|
||||
return 1;
|
||||
}
|
||||
|
||||
let schema;
|
||||
try {
|
||||
schema = loadSchema(path.join(__dirname, "report-schema.json"));
|
||||
} catch (error) {
|
||||
console.error("Failed to load report-schema.json:", error.message);
|
||||
return 1;
|
||||
}
|
||||
|
||||
let contents;
|
||||
try {
|
||||
contents = readFileWithinLimit(file);
|
||||
} catch (error) {
|
||||
const reason = error instanceof SafeInputError ? error.message : "input could not be opened or read safely";
|
||||
console.error(`Failed to read findings JSON: ${reason}`);
|
||||
return 1;
|
||||
}
|
||||
|
||||
try {
|
||||
enforceJsonTextLimits(contents);
|
||||
} catch (error) {
|
||||
const reason = error instanceof JsonStructureError ? error.message : "invalid JSON structure";
|
||||
console.error(`Failed to parse findings JSON: ${reason}`);
|
||||
return 1;
|
||||
}
|
||||
|
||||
let findings;
|
||||
try {
|
||||
findings = JSON.parse(contents);
|
||||
} catch {
|
||||
console.error("Failed to parse findings JSON: invalid JSON syntax");
|
||||
return 1;
|
||||
}
|
||||
|
||||
let errors;
|
||||
try {
|
||||
errors = validateDocument(findings, schema);
|
||||
} catch {
|
||||
console.error("Failed to validate findings JSON: unexpected validation error");
|
||||
return 1;
|
||||
}
|
||||
for (const message of errors) console.error("ERROR:", escapeUnsafeDiagnosticCharacters(message));
|
||||
if (errors.length > 0) {
|
||||
const cap = errors.length === LIMITS.validationErrors ? `; output capped at ${LIMITS.validationErrors}` : "";
|
||||
console.error(`FAIL: ${errors.length} validation error(s)${cap}`);
|
||||
return 1;
|
||||
}
|
||||
console.log(`PASS: ${findings.length} findings valid`);
|
||||
return 0;
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
LIMITS,
|
||||
PATH_FORBIDDEN_CHARACTER,
|
||||
UNSAFE_DIAGNOSTIC_CHARACTER,
|
||||
VISIBLE_CONTENT,
|
||||
WINDOWS_RESERVED_COMPONENT,
|
||||
collect,
|
||||
collectFindingSemanticErrors,
|
||||
collectSchemaErrors,
|
||||
hasVisibleProse,
|
||||
isSafeRelativeSourcePath,
|
||||
validateDocument,
|
||||
};
|
||||
|
||||
if (require.main === module) process.exit(run(process.argv[2]));
|
||||
652
.agents/skills/security-audit/validate-findings.test.cjs
Normal file
652
.agents/skills/security-audit/validate-findings.test.cjs
Normal file
@@ -0,0 +1,652 @@
|
||||
const assert = require("node:assert/strict");
|
||||
const fs = require("node:fs");
|
||||
const os = require("node:os");
|
||||
const path = require("node:path");
|
||||
const { spawnSync } = require("node:child_process");
|
||||
const test = require("node:test");
|
||||
const schema = require("./report-schema.json");
|
||||
const {
|
||||
LIMITS,
|
||||
collect,
|
||||
collectSchemaErrors,
|
||||
validateDocument,
|
||||
} = require("./validate-findings.cjs");
|
||||
|
||||
const validatorPath = path.join(__dirname, "validate-findings.cjs");
|
||||
const CLI_TIMEOUT_MS = 5000;
|
||||
const HOSTILE_CLI_TIMEOUT_MS = 15000;
|
||||
const HAS_SAFE_INPUT_OPEN = Number.isInteger(fs.constants.O_NOFOLLOW) &&
|
||||
fs.constants.O_NOFOLLOW !== 0 &&
|
||||
Number.isInteger(fs.constants.O_NONBLOCK) &&
|
||||
fs.constants.O_NONBLOCK !== 0;
|
||||
const TERMINAL_CONTROL_PAYLOAD = "\u001b\u0007\u0085\u202e\u034f\ufe0f";
|
||||
const TERMINAL_CONTROL_BYTES = [
|
||||
Buffer.from([0x1b]),
|
||||
Buffer.from([0x07]),
|
||||
Buffer.from("\u0085"),
|
||||
Buffer.from("\u202e"),
|
||||
Buffer.from("\u034f"),
|
||||
Buffer.from("\ufe0f"),
|
||||
];
|
||||
|
||||
function source(kind = "entrypoint", file = "src/handler.c", line = 10) {
|
||||
return { kind, file, line, scope: "handle", description: "Attacker data reaches the operation." };
|
||||
}
|
||||
|
||||
function evidence(file = "src/handler.c", line = 10) {
|
||||
return { file, line, description: "The source performs the operation without the required check." };
|
||||
}
|
||||
|
||||
function confirmed() {
|
||||
return {
|
||||
verdict: "confirmed",
|
||||
fingerprint: "src-handler-missing-check",
|
||||
title: "Missing ownership check",
|
||||
description: "An attacker can reach an operation without the intended ownership check.",
|
||||
root_cause: "handle omits the ownership check before changing the object.",
|
||||
intended_behavior: "Only the object's owner can change it.",
|
||||
trace: [source("entrypoint"), source("propagation", "src/model.c", 20), source("sink", "src/store.c", 30)],
|
||||
evidence: [evidence()],
|
||||
conditions: [],
|
||||
execution: {
|
||||
attacker_perspective: "An unprivileged remote user with their own account.",
|
||||
payloads: ["An object identifier owned by another user."],
|
||||
instructions: ["Submit the identifier through the public operation."],
|
||||
observed_result: "The other user's object changes.",
|
||||
},
|
||||
remediation: { strategy: "Check ownership before the state change." },
|
||||
severity: {
|
||||
likelihood: { score: "medium", reason: "The operation is directly reachable." },
|
||||
impact: { score: "medium", reason: "The attacker changes one protected object." },
|
||||
overall_severity: "medium",
|
||||
},
|
||||
confidence: { score: "high", reason: "The source path and result were reproduced." },
|
||||
};
|
||||
}
|
||||
|
||||
function needsValidation() {
|
||||
return {
|
||||
verdict: "needs_validation",
|
||||
fingerprint: "src-parser-size-hypothesis",
|
||||
title: "Unchecked parsed size",
|
||||
description: "A parsed size may reach an allocation without a limit.",
|
||||
claimed_root_cause: "parse_size may pass an unbounded value to allocate.",
|
||||
trace: [source()],
|
||||
evidence: [evidence()],
|
||||
blockers: ["The generated parser source is absent from this checkout."],
|
||||
validation_plan: {
|
||||
local: "Generate the parser and submit the smallest input that exceeds the documented limit.",
|
||||
deployment: "In an approved test deployment, confirm the request reaches the generated parser and record the bounded observable result.",
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
function rejected() {
|
||||
return {
|
||||
verdict: "rejected",
|
||||
fingerprint: "src-router-auth-bypass",
|
||||
title: "Authorization bypass in router",
|
||||
description: "The candidate claimed a route bypassed authorization.",
|
||||
claimed_root_cause: "dispatch was claimed to skip the authorization wrapper.",
|
||||
trace: [source("sink")],
|
||||
evidence: [evidence()],
|
||||
reason: "All routes pass through the authorization wrapper before dispatch.",
|
||||
};
|
||||
}
|
||||
|
||||
function errorsFor(value) {
|
||||
return validateDocument(value, schema);
|
||||
}
|
||||
|
||||
function runCli(contents, options = {}) {
|
||||
const { nodeArgs = [], timeout = CLI_TIMEOUT_MS } = options;
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-findings-"));
|
||||
const findingsPath = path.join(directory, "findings.json");
|
||||
try {
|
||||
fs.writeFileSync(findingsPath, contents);
|
||||
return spawnSync(process.execPath, [...nodeArgs, validatorPath, findingsPath], {
|
||||
encoding: "utf8",
|
||||
timeout,
|
||||
});
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
}
|
||||
|
||||
function cliOutput(result) {
|
||||
return `${result.stdout}${result.stderr}`;
|
||||
}
|
||||
|
||||
function assertNoInjectedControlBytes(output) {
|
||||
const bytes = Buffer.isBuffer(output) ? output : Buffer.from(output, "utf8");
|
||||
for (const marker of TERMINAL_CONTROL_BYTES) {
|
||||
assert.equal(bytes.indexOf(marker), -1, `found raw control bytes ${marker.toString("hex")}`);
|
||||
}
|
||||
}
|
||||
|
||||
function producerShapedFindings() {
|
||||
const demonstrated = confirmed();
|
||||
demonstrated.conditions = [{
|
||||
kind: "authentication_level",
|
||||
description: "The attacker needs a normal account.",
|
||||
}];
|
||||
demonstrated.execution.payloads = ["", " \t\r\n", "\u0000\u001f\u007f", "\u034f", "\ufe0f", "\ud800", "\udc00", "[{,}]\\\""];
|
||||
demonstrated.remediation.code_changes = [{
|
||||
file_name: "src/handler.c",
|
||||
fixed_code: "",
|
||||
}];
|
||||
|
||||
const blocked = needsValidation();
|
||||
delete blocked.validation_plan.deployment;
|
||||
|
||||
return [demonstrated, blocked, rejected()];
|
||||
}
|
||||
|
||||
function rejectMutation(factory, mutate) {
|
||||
const value = factory();
|
||||
mutate(value);
|
||||
assert.notEqual(errorsFor([value]).length, 0);
|
||||
}
|
||||
|
||||
test("schema is an actual top-level array with exactly three branches", () => {
|
||||
assert.equal(schema.type, "array");
|
||||
assert.equal(schema.items.oneOf.length, 3);
|
||||
assert.deepEqual(schema.items.oneOf.map((branch) => branch.properties.verdict.const), [
|
||||
"confirmed", "needs_validation", "rejected",
|
||||
]);
|
||||
const confirmedSchema = schema.items.oneOf[0].properties;
|
||||
assert.equal(confirmedSchema.title.visibleContent, true);
|
||||
assert.equal(confirmedSchema.execution.properties.payloads.items.minLength, undefined);
|
||||
assert.equal(confirmedSchema.execution.properties.payloads.items.visibleContent, undefined);
|
||||
assert.equal(confirmedSchema.remediation.properties.code_changes.items.properties.fixed_code.minLength, undefined);
|
||||
});
|
||||
|
||||
test("accepts a producer-shaped findings document through the CLI", () => {
|
||||
const result = runCli(JSON.stringify(producerShapedFindings()));
|
||||
assert.equal(result.status, 0, cliOutput(result));
|
||||
assert.match(result.stdout, /PASS: 3 findings valid/);
|
||||
});
|
||||
|
||||
test("accepts empty output and each complete branch", () => {
|
||||
assert.deepEqual(errorsFor([]), []);
|
||||
assert.deepEqual(errorsFor([confirmed(), needsValidation(), rejected()]), []);
|
||||
const localOnly = needsValidation();
|
||||
delete localOnly.validation_plan.deployment;
|
||||
assert.deepEqual(errorsFor([localOnly]), []);
|
||||
const deploymentOnly = needsValidation();
|
||||
delete deploymentOnly.validation_plan.local;
|
||||
assert.deepEqual(errorsFor([deploymentOnly]), []);
|
||||
});
|
||||
|
||||
test("allows a one-line finding trace", () => {
|
||||
const finding = confirmed();
|
||||
finding.trace = [source("entrypoint")];
|
||||
assert.deepEqual(errorsFor([finding]), []);
|
||||
});
|
||||
|
||||
test("rejects empty required content", () => {
|
||||
const cases = [
|
||||
[confirmed, (finding) => { finding.title = ""; }],
|
||||
[confirmed, (finding) => { finding.title = " "; }],
|
||||
[confirmed, (finding) => { finding.evidence = []; }],
|
||||
[confirmed, (finding) => { finding.execution.payloads = []; }],
|
||||
[confirmed, (finding) => { finding.execution.instructions = []; }],
|
||||
[confirmed, (finding) => { finding.execution.observed_result = ""; }],
|
||||
[confirmed, (finding) => { finding.remediation.strategy = ""; }],
|
||||
[needsValidation, (finding) => { finding.blockers = []; }],
|
||||
[needsValidation, (finding) => { finding.validation_plan = {}; }],
|
||||
[needsValidation, (finding) => { finding.validation_plan = { local: " " }; }],
|
||||
[rejected, (finding) => { finding.claimed_root_cause = ""; }],
|
||||
];
|
||||
for (const [factory, mutate] of cases) rejectMutation(factory, mutate);
|
||||
});
|
||||
|
||||
test("preserves exact payload and replacement-code strings", () => {
|
||||
const finding = confirmed();
|
||||
const payloads = ["", " \t\r\n", "\u0000\u001f\u007f", "\u034f", "\ufe0f", "\ud800", "\udc00"];
|
||||
const fixedCode = "\u0000 \t\r\n\u001f\u007f\u034f\ufe0f\ud800x\udc00";
|
||||
finding.execution.payloads = payloads.slice();
|
||||
finding.remediation.code_changes = [{ file_name: "src/handler.c", fixed_code: fixedCode }];
|
||||
|
||||
assert.deepEqual(errorsFor([finding]), []);
|
||||
assert.deepEqual(finding.execution.payloads, payloads);
|
||||
assert.equal(finding.remediation.code_changes[0].fixed_code, fixedCode);
|
||||
});
|
||||
|
||||
test("rejects invalid scalars and whitespace, control, format, or default-ignorable prose", () => {
|
||||
for (const invisible of ["\u0000\t\r\n\u001f\u007f\u200b", "\u034f", "\ufe0f", "\ud800", "\udc00", "visible\ud800"]) {
|
||||
const cases = [
|
||||
[confirmed, (finding) => { finding.title = invisible; }],
|
||||
[confirmed, (finding) => { finding.trace[0].scope = invisible; }],
|
||||
[confirmed, (finding) => { finding.evidence[0].description = invisible; }],
|
||||
[confirmed, (finding) => { finding.execution.instructions = [invisible]; }],
|
||||
[confirmed, (finding) => { finding.remediation.strategy = invisible; }],
|
||||
[confirmed, (finding) => { finding.severity.impact.reason = invisible; }],
|
||||
[confirmed, (finding) => { finding.confidence.reason = invisible; }],
|
||||
[needsValidation, (finding) => { finding.blockers = [invisible]; }],
|
||||
[needsValidation, (finding) => { finding.validation_plan = { local: invisible }; }],
|
||||
[rejected, (finding) => { finding.reason = invisible; }],
|
||||
];
|
||||
for (const [factory, mutate] of cases) rejectMutation(factory, mutate);
|
||||
}
|
||||
});
|
||||
|
||||
test("quotes input-derived controls in direct validation values and paths", () => {
|
||||
const finding = confirmed();
|
||||
finding.trace[0].kind = `invalid-${TERMINAL_CONTROL_PAYLOAD}`;
|
||||
finding.execution[`extra-${TERMINAL_CONTROL_PAYLOAD}`] = "value";
|
||||
|
||||
const cyclic = {};
|
||||
cyclic[`path-${TERMINAL_CONTROL_PAYLOAD}`] = cyclic;
|
||||
const output = [
|
||||
...errorsFor([finding]),
|
||||
...collect(cyclic, { type: "object" }, "$input"),
|
||||
].join("\n");
|
||||
|
||||
for (const escaped of ["\\u001b", "\\u0007", "\\u0085", "\\u202e", "\\u034f", "\\ufe0f"]) {
|
||||
assert(output.includes(escaped), `missing escaped diagnostic ${escaped}`);
|
||||
}
|
||||
assertNoInjectedControlBytes(output);
|
||||
});
|
||||
|
||||
test("rejects line zero", () => {
|
||||
rejectMutation(confirmed, (finding) => { finding.trace[0].line = 0; });
|
||||
rejectMutation(rejected, (finding) => { finding.evidence[0].line = 0; });
|
||||
});
|
||||
|
||||
test("does not treat inherited or Object-prototype properties as schema properties", () => {
|
||||
rejectMutation(confirmed, (finding) => { finding.constructor = "not allowed"; });
|
||||
const inherited = Object.create({ verdict: "confirmed" });
|
||||
assert(errorsFor([inherited]).some((error) => error.includes("exactly one")));
|
||||
assert(collect(Object.create({ constructor: "inherited" }), {
|
||||
type: "object",
|
||||
properties: { constructor: { type: "string" } },
|
||||
required: ["constructor"],
|
||||
additionalProperties: false,
|
||||
}).some((error) => error.includes("missing required")));
|
||||
});
|
||||
|
||||
test("oneOf requires exactly one passing branch", () => {
|
||||
assert(collect("value", { oneOf: [{ type: "string" }, { minLength: 1 }] }, "$test")
|
||||
.some((error) => error.includes("matched 2")));
|
||||
assert(collect(7, { oneOf: [{ type: "string" }, { minimum: 10 }] }, "$test")
|
||||
.some((error) => error.includes("matched 0")));
|
||||
});
|
||||
|
||||
test("rejects duplicate fingerprints and unique array entries", () => {
|
||||
const first = confirmed();
|
||||
const second = rejected();
|
||||
second.fingerprint = first.fingerprint;
|
||||
assert(errorsFor([first, second]).some((error) => error.includes("duplicate of")));
|
||||
rejectMutation(needsValidation, (finding) => { finding.blockers = [finding.blockers[0], finding.blockers[0]]; });
|
||||
});
|
||||
|
||||
test("uses canonical Set uniqueness for structured entries at the array limit", () => {
|
||||
const entries = Array.from({ length: LIMITS.arrayItems }, (_, id) => ({ id, label: String(id) }));
|
||||
assert.deepEqual(collect(entries, { type: "array", uniqueItems: true }), []);
|
||||
|
||||
const duplicate = entries.slice(0, -1);
|
||||
duplicate.push({ label: "0", id: 0 });
|
||||
assert(collect(duplicate, { type: "array", uniqueItems: true })
|
||||
.some((error) => error.includes(`duplicate at index ${LIMITS.arrayItems - 1}`)));
|
||||
});
|
||||
|
||||
test("bounds canonical uniqueness keys and Set storage", () => {
|
||||
const oversizedKey = "x".repeat(LIMITS.canonicalKeyBytes + 1);
|
||||
assert(collect([oversizedKey], { type: "array", uniqueItems: true })
|
||||
.some((error) => error.includes("canonical key exceeds")));
|
||||
|
||||
const itemLength = Math.floor(LIMITS.uniqueSetBytes / 6);
|
||||
const largeUniqueItems = Array.from({ length: 6 }, (_, index) => `${index}${"x".repeat(itemLength)}`);
|
||||
assert(collect(largeUniqueItems, { type: "array", uniqueItems: true })
|
||||
.some((error) => error.includes("canonical uniqueness set exceeds")));
|
||||
});
|
||||
|
||||
test("requires findings to be sorted by fingerprint", () => {
|
||||
const first = confirmed();
|
||||
const second = rejected();
|
||||
assert(errorsFor([second, first]).some((error) => error.includes("sorted lexicographically")));
|
||||
});
|
||||
|
||||
test("rejects severity above demonstrated impact", () => {
|
||||
rejectMutation(confirmed, (finding) => {
|
||||
finding.severity.overall_severity = "high";
|
||||
finding.severity.impact.score = "medium";
|
||||
});
|
||||
});
|
||||
|
||||
test("rejects unsafe source paths", () => {
|
||||
const badPaths = [
|
||||
"/etc/passwd",
|
||||
"../src/file.c",
|
||||
"src/../file.c",
|
||||
"src//file.c",
|
||||
"C:\\src\\file.c",
|
||||
"src/file:name.c",
|
||||
"src/file\nname.c",
|
||||
"src/file\u0001name.c",
|
||||
"src/file\u0085name.c",
|
||||
"src/file\u2028name.c",
|
||||
"src/file\u202ename.c",
|
||||
"src/file\u2066name.c",
|
||||
"src/file\u200dname.c",
|
||||
"src/file\u034fname.c",
|
||||
"src/file\ufe0fname.c",
|
||||
"src/file\ud800name.c",
|
||||
"src/file\udc00name.c",
|
||||
"CON",
|
||||
"src/con.txt",
|
||||
"src/PRN",
|
||||
"src/AUX.c",
|
||||
"src/NUL",
|
||||
"src/COM1.log",
|
||||
"src/lpt9",
|
||||
"src/CONIN$",
|
||||
"src/CONOUT$.txt",
|
||||
"src/CLOCK$.txt",
|
||||
"src/COM\u00b9.log",
|
||||
"src/LPT\u00b2.log",
|
||||
"src /file.c",
|
||||
"src./file.c",
|
||||
"src/file.c ",
|
||||
"src/file.c.",
|
||||
];
|
||||
for (const badPath of badPaths) {
|
||||
rejectMutation(confirmed, (finding) => { finding.trace[0].file = badPath; });
|
||||
}
|
||||
rejectMutation(rejected, (finding) => { finding.evidence[0].file = "NUL.txt"; });
|
||||
rejectMutation(confirmed, (finding) => {
|
||||
finding.remediation.code_changes = [{ file_name: "src/file:name.c", fixed_code: "replacement" }];
|
||||
});
|
||||
});
|
||||
|
||||
test("accepts legitimate Unicode source paths and prose", () => {
|
||||
const finding = confirmed();
|
||||
finding.title = "Finding \ud83d\ude00 cafe\u0301";
|
||||
finding.trace[0].file = "src/日本語/cafe\u0301-\ud83d\ude00.ts";
|
||||
finding.evidence[0].file = "src/mañana/файл.ts";
|
||||
finding.remediation.code_changes = [{
|
||||
file_name: "src/修正/éxito.ts",
|
||||
fixed_code: "replacement",
|
||||
}];
|
||||
assert.deepEqual(errorsFor([finding]), []);
|
||||
});
|
||||
|
||||
test("CLI rejects input above the byte limit without an exception trace", () => {
|
||||
const result = runCli(Buffer.alloc(LIMITS.inputBytes + 1, 0x20));
|
||||
const output = cliOutput(result);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, new RegExp(`input exceeds ${LIMITS.inputBytes} byte limit`));
|
||||
assert.doesNotMatch(output, /RangeError|Maximum call stack|heap out of memory/i);
|
||||
});
|
||||
|
||||
test("CLI rejects invalid UTF-8 without replacement or an exception trace", () => {
|
||||
const findings = producerShapedFindings();
|
||||
findings[0].execution.payloads = ["INVALID_UTF8"];
|
||||
const encoded = Buffer.from(JSON.stringify(findings));
|
||||
const marker = Buffer.from("INVALID_UTF8");
|
||||
const markerOffset = encoded.indexOf(marker);
|
||||
assert.notEqual(markerOffset, -1);
|
||||
const malformed = Buffer.concat([
|
||||
encoded.subarray(0, markerOffset),
|
||||
Buffer.from([0x80]),
|
||||
encoded.subarray(markerOffset + marker.length),
|
||||
]);
|
||||
|
||||
const result = runCli(malformed);
|
||||
const output = cliOutput(result);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /input is not valid UTF-8/);
|
||||
assert.doesNotMatch(output, /TypeError|stack|at validate-findings/i);
|
||||
});
|
||||
|
||||
test("quotes input-derived controls in CLI validation errors", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const finding = confirmed();
|
||||
finding.trace[0].kind = `invalid-${TERMINAL_CONTROL_PAYLOAD}`;
|
||||
finding.execution[`extra-${TERMINAL_CONTROL_PAYLOAD}`] = "value";
|
||||
const result = runCli(JSON.stringify([finding]));
|
||||
|
||||
assert.equal(result.status, 1, cliOutput(result));
|
||||
assert.match(result.stderr, /\$\[0\]\.trace\[0\]\.kind/);
|
||||
assert.match(result.stderr, /\\u001b/);
|
||||
assert.match(result.stderr, /\\u202e/);
|
||||
assertNoInjectedControlBytes(result.stderr);
|
||||
});
|
||||
|
||||
test("returns a generic syntax error without parser-supplied controls", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const malformed = Buffer.concat([
|
||||
Buffer.from("["),
|
||||
Buffer.from(TERMINAL_CONTROL_PAYLOAD),
|
||||
Buffer.from("]"),
|
||||
]);
|
||||
const result = runCli(malformed);
|
||||
|
||||
assert.equal(result.status, 1, cliOutput(result));
|
||||
assert.equal(result.stderr, "Failed to parse findings JSON: invalid JSON syntax\n");
|
||||
assertNoInjectedControlBytes(result.stderr);
|
||||
});
|
||||
|
||||
test("does not reflect controls from a failed CLI input path", () => {
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-findings-path-"));
|
||||
const missingPath = path.join(directory, `missing-${TERMINAL_CONTROL_PAYLOAD}.json`);
|
||||
try {
|
||||
const result = spawnSync(process.execPath, [validatorPath, missingPath], {
|
||||
encoding: "utf8",
|
||||
timeout: CLI_TIMEOUT_MS,
|
||||
});
|
||||
assert.equal(result.status, 1, cliOutput(result));
|
||||
assert.match(result.stderr, /Failed to read findings JSON:/);
|
||||
assertNoInjectedControlBytes(result.stderr);
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("CLI rejects lone-surrogate prose without changing payload semantics", () => {
|
||||
const findings = producerShapedFindings();
|
||||
findings[0].title = "\ud800";
|
||||
const result = runCli(JSON.stringify(findings));
|
||||
const output = cliOutput(result);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /must contain only valid Unicode scalar values/);
|
||||
assert.doesNotMatch(output, /stack|at validate-findings/i);
|
||||
});
|
||||
|
||||
test("CLI rejects Unicode format controls in source paths", () => {
|
||||
const findings = producerShapedFindings();
|
||||
findings[0].trace[0].file = "src/file\u202ename.c";
|
||||
const result = runCli(JSON.stringify(findings));
|
||||
const output = cliOutput(result);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /must be a safe repository-relative source path/);
|
||||
assert.doesNotMatch(output, /stack|at validate-findings/i);
|
||||
});
|
||||
|
||||
test("CLI rejects a FIFO without blocking", { skip: process.platform === "win32" || !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-findings-fifo-"));
|
||||
const fifoPath = path.join(directory, "findings.json");
|
||||
try {
|
||||
const created = spawnSync("mkfifo", [fifoPath], { encoding: "utf8", timeout: CLI_TIMEOUT_MS });
|
||||
assert.equal(created.status, 0, cliOutput(created));
|
||||
|
||||
const result = spawnSync(process.execPath, [validatorPath, fifoPath], {
|
||||
encoding: "utf8",
|
||||
timeout: CLI_TIMEOUT_MS,
|
||||
});
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /input must be a regular file/);
|
||||
assert.doesNotMatch(output, /stack|at validate-findings/i);
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("CLI rejects a symlink without following it", { skip: process.platform === "win32" || !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), "validate-findings-symlink-"));
|
||||
const targetPath = path.join(directory, "target.json");
|
||||
const symlinkPath = path.join(directory, "findings.json");
|
||||
try {
|
||||
fs.writeFileSync(targetPath, JSON.stringify(producerShapedFindings()));
|
||||
fs.symlinkSync(targetPath, symlinkPath);
|
||||
const result = spawnSync(process.execPath, [validatorPath, symlinkPath], {
|
||||
encoding: "utf8",
|
||||
timeout: CLI_TIMEOUT_MS,
|
||||
});
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /input must not be a symlink/);
|
||||
assert.doesNotMatch(output, /stack|at validate-findings/i);
|
||||
} finally {
|
||||
fs.rmSync(directory, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("CLI rejects input above the nesting-depth limit without an exception trace", () => {
|
||||
const levels = LIMITS.nestingDepth + 1;
|
||||
const result = runCli(`${"[".repeat(levels)}0${"]".repeat(levels)}`);
|
||||
const output = cliOutput(result);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, new RegExp(`${LIMITS.nestingDepth} level nesting depth limit`));
|
||||
assert.doesNotMatch(output, /RangeError|Maximum call stack|heap out of memory/i);
|
||||
});
|
||||
|
||||
test("CLI rejects an oversized array without an exception trace", () => {
|
||||
const result = runCli(JSON.stringify(Array(LIMITS.arrayItems + 1).fill(null)));
|
||||
const output = cliOutput(result);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, new RegExp(`${LIMITS.arrayItems} item array limit`));
|
||||
assert.doesNotMatch(output, /RangeError|Maximum call stack|heap out of memory/i);
|
||||
});
|
||||
|
||||
test("checks pattern and branch invariants", () => {
|
||||
rejectMutation(rejected, (finding) => { finding.fingerprint = "not stable"; });
|
||||
rejectMutation(needsValidation, (finding) => {
|
||||
finding.severity = { impact: { score: "low" }, overall_severity: "low" };
|
||||
});
|
||||
});
|
||||
|
||||
test("rejects unsupported and malformed schema keywords", () => {
|
||||
assert(collectSchemaErrors({ type: "string", format: "uuid" }).some((error) => error.includes("format")));
|
||||
assert(collectSchemaErrors({ type: "string", pattern: "[" }).some((error) => error.includes("regular expression")));
|
||||
assert(collectSchemaErrors({ type: "string", visibleContent: "yes" }).some((error) => error.includes("expected boolean")));
|
||||
assert(collectSchemaErrors({ type: "array", visibleContent: true }).some((error) => error.includes("requires type")));
|
||||
assert.notEqual(validateDocument([], { type: "array", maxItems: 1 }).length, 0);
|
||||
});
|
||||
|
||||
test("caps malformed 1000-finding validation output", () => {
|
||||
assert.equal(errorsFor(Array.from({ length: LIMITS.arrayItems }, () => null)).length, LIMITS.validationErrors);
|
||||
if (!HAS_SAFE_INPUT_OPEN) return;
|
||||
|
||||
const result = runCli(JSON.stringify(Array.from({ length: LIMITS.arrayItems }, () => null)));
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /output capped at 100/);
|
||||
assert(output.length < 20000, `unexpected output length ${output.length}`);
|
||||
assert.doesNotMatch(output, /RangeError|Maximum call stack|stack|at validate-findings/i);
|
||||
});
|
||||
|
||||
test("caps amplified in-limit findings output under a constrained Node heap", { skip: !HAS_SAFE_INPUT_OPEN }, () => {
|
||||
const findings = Array.from({ length: 750 }, () => ({
|
||||
verdict: "confirmed",
|
||||
trace: Array.from({ length: LIMITS.arrayItems }, () => 0),
|
||||
evidence: Array.from({ length: LIMITS.arrayItems }, () => 0),
|
||||
}));
|
||||
const contents = JSON.stringify(findings);
|
||||
assert(contents.length > 3 * 1000 * 1000, `hostile input too small: ${contents.length}`);
|
||||
assert(contents.length <= LIMITS.inputBytes, `hostile input over limit: ${contents.length}`);
|
||||
|
||||
const result = runCli(contents, {
|
||||
nodeArgs: ["--max-old-space-size=64"],
|
||||
timeout: HOSTILE_CLI_TIMEOUT_MS,
|
||||
});
|
||||
const output = cliOutput(result);
|
||||
assert.notEqual(result.error && result.error.code, "ETIMEDOUT", output);
|
||||
assert.equal(result.status, 1, output);
|
||||
assert.match(output, /output capped at 100/);
|
||||
assert(output.length < 20000, `unexpected output length ${output.length}`);
|
||||
assert.doesNotMatch(output, /heap out of memory|allocation failed|RangeError|Maximum call stack/i);
|
||||
});
|
||||
|
||||
test("keeps shared helpers aligned with the coverage-ledger validator", () => {
|
||||
const findingsModule = require("./validate-findings.cjs");
|
||||
const ledgerModule = require("./validate-coverage-ledger.cjs");
|
||||
|
||||
for (const name of [
|
||||
"VISIBLE_CONTENT",
|
||||
"PATH_FORBIDDEN_CHARACTER",
|
||||
"WINDOWS_RESERVED_COMPONENT",
|
||||
"UNSAFE_DIAGNOSTIC_CHARACTER",
|
||||
]) {
|
||||
assert.equal(findingsModule[name].source, ledgerModule[name].source, `${name} source`);
|
||||
assert.equal(findingsModule[name].flags, ledgerModule[name].flags, `${name} flags`);
|
||||
}
|
||||
|
||||
const sharedLimitKeys = Object.keys(findingsModule.LIMITS)
|
||||
.filter((key) => Object.prototype.hasOwnProperty.call(ledgerModule.LIMITS, key))
|
||||
.sort();
|
||||
assert.deepEqual(sharedLimitKeys, ["inputBytes", "nestingDepth", "validationErrors"]);
|
||||
for (const key of sharedLimitKeys) {
|
||||
assert.equal(findingsModule.LIMITS[key], ledgerModule.LIMITS[key], `LIMITS.${key}`);
|
||||
}
|
||||
|
||||
const pathCorpus = [
|
||||
"src/handler.js",
|
||||
"src/caf\u00e9/handler.js",
|
||||
"src/\u65e5\u672c\u8a9e/\u0444\u0430\u0439\u043b.ts",
|
||||
"src/cloc\u212a$.txt",
|
||||
"src/CLOCK$.txt",
|
||||
"src/con.txt",
|
||||
"CON",
|
||||
"src/COM\u00b9.log",
|
||||
"src/lpt\u00b3",
|
||||
"/etc/passwd",
|
||||
"../src/file.c",
|
||||
"src/../file.c",
|
||||
"src//file.c",
|
||||
"src\\file.c",
|
||||
"src/file:name.c",
|
||||
"~home/file.c",
|
||||
"C:/file.c",
|
||||
"src/file.c ",
|
||||
"src/file.c.",
|
||||
"src/file\u202ename.c",
|
||||
"src/file\u200b.js",
|
||||
"src/file\u034f.js",
|
||||
"src/file\ufe0f.js",
|
||||
"src/file\ud800name.c",
|
||||
"src/file\udc00name.c",
|
||||
];
|
||||
for (const value of pathCorpus) {
|
||||
assert.equal(
|
||||
findingsModule.isSafeRelativeSourcePath(value),
|
||||
ledgerModule.isSafeRelativePath(value),
|
||||
`path verdict diverges for ${JSON.stringify(value)}`,
|
||||
);
|
||||
}
|
||||
assert.equal(findingsModule.isSafeRelativeSourcePath("src/cloc\u212a$.txt"), false);
|
||||
assert.equal(ledgerModule.isSafeRelativePath("src/cloc\u212a$.txt"), false);
|
||||
|
||||
const proseCorpus = [
|
||||
"Valid prose.",
|
||||
"caf\u00e9",
|
||||
"",
|
||||
" \t\r\n",
|
||||
"\u200b",
|
||||
"\u034f",
|
||||
"\ufe0f",
|
||||
"\ud800",
|
||||
"\udc00",
|
||||
"visible\ud800",
|
||||
];
|
||||
for (const value of proseCorpus) {
|
||||
assert.equal(
|
||||
findingsModule.hasVisibleProse(value),
|
||||
ledgerModule.hasVisibleProse(value),
|
||||
`prose verdict diverges for ${JSON.stringify(value)}`,
|
||||
);
|
||||
}
|
||||
});
|
||||
9
.claude/settings.json
Normal file
9
.claude/settings.json
Normal file
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"permissions": {
|
||||
"allow": [
|
||||
"Bash(ffprobe -v error *)",
|
||||
"Bash(env)",
|
||||
"mcp__outline__read_document"
|
||||
]
|
||||
}
|
||||
}
|
||||
14
.claude/settings.local.json
Normal file
14
.claude/settings.local.json
Normal file
@@ -0,0 +1,14 @@
|
||||
{
|
||||
"permissions": {
|
||||
"allow": [
|
||||
"Bash(rtk grep *)",
|
||||
"Bash(rtk read *)",
|
||||
"Bash(rtk git *)"
|
||||
],
|
||||
"additionalDirectories": [
|
||||
"/config/.claude/skills/security-audit",
|
||||
"/config/security-audit-skill",
|
||||
"/config/.cargo/registry"
|
||||
]
|
||||
}
|
||||
}
|
||||
1
.claude/skills/security-audit
Symbolic link
1
.claude/skills/security-audit
Symbolic link
@@ -0,0 +1 @@
|
||||
../../.agents/skills/security-audit
|
||||
13
.rtk/filters.toml
Normal file
13
.rtk/filters.toml
Normal file
@@ -0,0 +1,13 @@
|
||||
# Project-local RTK filters — commit this file with your repo.
|
||||
# Filters here override user-global and built-in filters.
|
||||
# Docs: https://github.com/rtk-ai/rtk#custom-filters
|
||||
schema_version = 1
|
||||
|
||||
# Example: suppress build noise from a custom tool
|
||||
# [filters.my-tool]
|
||||
# description = "Compact my-tool output"
|
||||
# match_command = "^my-tool\\s+build"
|
||||
# strip_ansi = true
|
||||
# strip_lines_matching = ["^\\s*$", "^Downloading", "^Installing"]
|
||||
# max_lines = 30
|
||||
# on_empty = "my-tool: ok"
|
||||
231
CHANGELOG.md
231
CHANGELOG.md
@@ -10,6 +10,228 @@ The long form, with what was wrong before and how it was found, is in
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
## [0.6.0] - 2026-09-15
|
||||
|
||||
### Added
|
||||
|
||||
- Each episode in Currently Listening has a cross that takes it off the list. It forgets where you
|
||||
got to, so playing it again starts from the beginning.
|
||||
- An admin can give a feed a Directory category in its settings (`category` in config.toml), for
|
||||
the blogs and other feeds that name none of their own. A feed's own iTunes category still wins.
|
||||
- Keyboard shortcuts after Feedly's: j and k through items, Shift-J and Shift-K through feeds,
|
||||
g and a letter to go to a place, o to play, s to pin, and more. Press ? for the whole list.
|
||||
- Directory can be filtered to Podcasts or Blogs, and by each show's own iTunes category as a row
|
||||
of chips, the narrower one where a show gives two (Games, not Leisure). The two combine, and
|
||||
both filter in place.
|
||||
|
||||
### Changed
|
||||
|
||||
- Currently Listening marks the episode in the player with the EQ bars, as the item list does,
|
||||
and its progress and time left move as it plays. Each row says how much is left.
|
||||
- Directory shows each feed as its cover art in a grid, title and subscriber count underneath,
|
||||
instead of a list. Popular and the Add a feed dialog keep their rows.
|
||||
- The pages are set in Inter, served by ipx itself. Classic keeps Lucida Grande.
|
||||
- Keeping an item is now pinning it: a thumbtack in place of the flag, and Pin, Pinned and Unpin
|
||||
in place of Keep, Kept and Stop keeping. A pinned item is still never deleted.
|
||||
- Currently Listening is its own place in the feed list, below Popular, instead of a section at
|
||||
the bottom of the Popular page.
|
||||
- The first scan after upgrading fetches every feed in full once, on its usual schedule, so each
|
||||
picks up its category without waiting for the publisher to change something.
|
||||
|
||||
### Fixed
|
||||
|
||||
- An episode that fails to load, or is paused before it has, no longer forgets where you left off
|
||||
in it, and so no longer drops out of Currently Listening.
|
||||
- Currently Listening lists the episodes you have started. It left out anything marked read, and
|
||||
opening an episode marks it read, so it usually showed nothing. An episode now leaves the list
|
||||
once 90% of it has played.
|
||||
- A WordPress post that embeds the file it encloses no longer lists, and downloads, that file
|
||||
twice. Items that already had it twice are folded into one at startup, and the spare copy
|
||||
deleted.
|
||||
- The pinned column's heading lines up with the pins under it, and every heading sits a pixel
|
||||
further right, over its column.
|
||||
- The feeds left behind by an OPML subscription removed before ipx retired them are cleared at
|
||||
startup: forgotten if nothing was downloaded, kept as orphaned if something was. Feeds from it
|
||||
that have since been given their own settings stay as they are, with their items. Removing an
|
||||
OPML or Patreon subscription no longer deletes the items of a feed inside it that has its own
|
||||
settings.
|
||||
- Titles that arrive as HTML, such as The Verge's, no longer show their entities as text:
|
||||
"Meta’s" reads "Meta’s". Titles already stored are corrected the next time their feed
|
||||
changes.
|
||||
|
||||
## [0.5.5] - 2026-09-14
|
||||
|
||||
### Changed
|
||||
|
||||
- A feed whose site sends a message instead of the feed, such as "Unable to establish a DB
|
||||
connection", now shows that message and is flagged as the publisher's problem, instead of two
|
||||
parser errors about reaching the end of input.
|
||||
|
||||
### Removed
|
||||
|
||||
- An unused icon glyph (`minus`) left over from before Unsubscribe settled on `circleMinus`.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Add a feed opened over Directory or Popular now shows its Popular list instead of staying on
|
||||
"Loading…", and no longer cuts the Directory behind it down to ten.
|
||||
|
||||
## [0.5.4] - 2026-09-14
|
||||
|
||||
### Added
|
||||
|
||||
- Currently Listening, below Popular: episodes you started and have not finished, across every
|
||||
feed you subscribe to. Tap one to pick up where you left off.
|
||||
- Theme has an Auto option, alongside Dark, Light and Classic, that follows your system's
|
||||
light/dark setting. All four are now also in Settings, as a dropdown next to the header
|
||||
button's one-click-at-a-time toggle -- the same setting either way.
|
||||
|
||||
### Changed
|
||||
|
||||
- The feed (or Directory/Popular/All Subscriptions) and the tab you had open are remembered
|
||||
across a reload or a new visit. A feed you no longer subscribe to, or a first visit with
|
||||
nothing remembered yet, lands on All Subscriptions instead of the first feed alphabetically.
|
||||
|
||||
### Fixed
|
||||
|
||||
- On iOS, the topbar (the hamburger menu included) could stop responding to taps until a hard
|
||||
refresh. The page sized itself with `100vh`, which iOS Safari measures against the address
|
||||
bar's collapsed state rather than what is actually visible; `100dvh` tracks the real viewport
|
||||
as the bar shows and hides.
|
||||
|
||||
## [0.5.3] - 2026-09-14
|
||||
|
||||
### Added
|
||||
|
||||
- A feed that has been failing for a day shows a plain-English reason in the sidebar and on its
|
||||
own page, sorted from a 404, a 401/403, a 402, a name that no longer resolves, or a web page in
|
||||
place of the feed -- with Unsubscribe or, when the page links its new feed, Use the new address.
|
||||
A feed that fails once and reads fine again within a day is never flagged.
|
||||
|
||||
### Changed
|
||||
|
||||
- Unsubscribing from the last person's OPML or Patreon subscription now retires the feeds it
|
||||
listed, the same as a feed the list itself drops: removed if nothing was downloaded, kept and
|
||||
marked orphaned otherwise. Until now they stayed in the database and kept being scanned hourly
|
||||
with auto-download on, which is how 922 defunct `davewiner` feeds outlived the OPML that listed
|
||||
them.
|
||||
|
||||
### Fixed
|
||||
|
||||
- A feed whose XML uses a bare `&` instead of `&` (kcpw, both feedland feeds) is now read
|
||||
instead of refused.
|
||||
- A feed URL that now serves a web page says so, and names the feed the page links to when it has
|
||||
one, instead of a raw XML parser error.
|
||||
- A publisher answering with an empty body (British Antarctic Survey's 202) is read as nothing new
|
||||
to report, not a parse failure.
|
||||
- A link in an item's show notes opens in a new tab instead of navigating away from ipx.
|
||||
- A video file plays as video, in a small floating pane above the player bar, instead of silently
|
||||
as sound only.
|
||||
- On the Unread tab, opening an item no longer makes it disappear from the list -- it stays until
|
||||
you open a different one, even if a scan finishes and refreshes the list while it is open.
|
||||
- Subscribe and Unsubscribe have their own icons (a circled check and a circled minus) instead of
|
||||
sharing the generic plus and minus used for adding feeds, users and imports.
|
||||
- Settings no longer disappears for a non-admin account. It was hiding the whole Settings modal
|
||||
along with the log and the users screen, but a non-admin has settings of their own in there --
|
||||
their subscriptions' Export and Import, and the schedule and quota are worth seeing even without
|
||||
a say in them. Only the log and the users screen, which the server also refuses them, are gone.
|
||||
|
||||
## [0.5.2] - 2026-09-12
|
||||
|
||||
### Added
|
||||
|
||||
- Settings → Users and `ipx user list` show when each account was added and when it last signed
|
||||
in, to the hour.
|
||||
- `ipx user rename <name> <new name>` renames an account and keeps its feeds, read state and admin
|
||||
rights. An account made before the proxy was set up can take the name the proxy signs it in as.
|
||||
|
||||
### Changed
|
||||
|
||||
- Directory and Popular list the feeds inside an OPML one by one, and no longer the OPML itself,
|
||||
so you can subscribe to just the shows you want.
|
||||
- The database no longer records when subscriptions and sign-in sessions were created. Nothing
|
||||
ever read it, and an existing database drops the columns on its next start.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Show notes that the podcast's host cut off in the middle of a tag no longer open with a scrap of
|
||||
HTML: the item's other copy of its notes is used instead, from the next time the feed changes.
|
||||
Daily Meditation Podcast had 57.
|
||||
- Docker no longer shows ipodderx as starting, or calls it unhealthy, while it scans or downloads:
|
||||
`ipx status` answers at once instead of waiting for the job in progress to finish.
|
||||
- Signing out after signing in through Cloudflare Access no longer lands on ipodderx's own password
|
||||
page. With the new `sign_out_url` set, Sign out ends the Access session, and the password page
|
||||
sends anyone the proxy signs in straight to their feeds.
|
||||
- The sign-in guide, `docs/sso.md`, describes the setup ipodderx.sdf1.net really runs: Authentik as
|
||||
Cloudflare Access's identity provider, and how to find the address ipx has to trust. It had never
|
||||
been checked against a real setup, and pointed at the wrong address.
|
||||
|
||||
## [0.5.1] - 2026-09-12
|
||||
|
||||
### Fixed
|
||||
|
||||
- The triangle that opens an OPML or Patreon folder was cramped against the folder's art. It has
|
||||
more room now, and a wider target to click.
|
||||
|
||||
## [0.5.0] - 2026-09-12
|
||||
|
||||
### Added
|
||||
|
||||
- A Patreon token pasted into Add feed, or a creator's RSS link without `&show=`, becomes a folder
|
||||
of that creator's shows, kept in step on every scan like a subscribed OPML. A creator with only
|
||||
one show stays a plain feed. One already added as a single long feed is split into its shows on
|
||||
its next scan, keeping its files and what you had read.
|
||||
- Add feed has an "Allow items marked explicit" box, so a new feed's first scan no longer skips
|
||||
every explicit item.
|
||||
|
||||
### Changed
|
||||
|
||||
- Unread counts, unread dots and download progress are amber, the colour of the icon's EQ bars.
|
||||
Blue is kept for the primary action and links, so a count no longer looks like a button.
|
||||
- What is playing is marked by small EQ bars, in its row and in the player. They move only while it
|
||||
plays.
|
||||
- A folder in the sidebar shows its first four shows' art as a mosaic, and its shows sit under its
|
||||
title. Only folders have a triangle, so every feed lines up with Directory and Popular above.
|
||||
- Feeds without art get initials in a colour of their own, instead of all the same grey.
|
||||
- The Flagged tab is called Kept, as the Keep button and Settings already said.
|
||||
- A feed's header is one short line; when it checks next is in its tooltip.
|
||||
- The item list takes more of the window, and the Files pane shows only when the item has files.
|
||||
- Column headings and tags are in sentence case, and fewer things are bold.
|
||||
- Unsubscribe is a round button beside the feed's other actions.
|
||||
- Export and Import in Settings say what they do.
|
||||
- The sign-in page shows the original icon large.
|
||||
- Nothing animates when your system asks for reduced motion.
|
||||
- The pages are about 90 KB smaller: the icon is served once instead of written into each.
|
||||
- A web token generated for a new install is 64 characters instead of 32.
|
||||
- The README is a short overview of what ipx does and how to run it, and points into `docs/` for
|
||||
the rest. It still described the layout from before 0.4.0.
|
||||
|
||||
### Removed
|
||||
|
||||
- The systemd units in `contrib/`. Run ipx with Docker, or point a unit of your own at
|
||||
`ipx daemon`.
|
||||
- Upgrading from before 0.3.0 directly: what was read, kept or part-played before accounts is no
|
||||
longer carried over to the admin, and OPML feeds that old versions wrote into `config.toml` are
|
||||
no longer moved out of it. Upgrade through 0.4.0 first.
|
||||
- `interval_mins` in `config.toml` is ignored; use `schedule`.
|
||||
|
||||
### Fixed
|
||||
|
||||
- The feed list works from the keyboard: Tab reaches every feed, Enter opens it, and Right and Left
|
||||
open and close a folder. Every button shows where the focus is, including in the toolbar, which
|
||||
used to clip the ring.
|
||||
- "1 items" reads "1 item".
|
||||
- A selected feed without art no longer loses its initials tile in Dark and Light.
|
||||
- The player shows the feed's initials when there is no art, not the episode's.
|
||||
- Turning on Allow explicit, or changing keywords or auto-download, brings back what those settings
|
||||
had skipped on the feed's next scan. Before, an item was judged once, when first seen, and a
|
||||
skipped one stayed skipped whatever you changed.
|
||||
- Feeds inside an OPML or a Patreon creator follow your settings on the folder unless you set their
|
||||
own, as the folder's settings dialog said they did. Before, the folder's settings reached nothing
|
||||
inside it.
|
||||
- A new feed no longer takes the name of one you removed earlier and shows that feed's old items.
|
||||
Re-adding the same feed still gets its old name, and its history, back.
|
||||
|
||||
## [0.4.0] - 2026-09-11
|
||||
|
||||
### Added
|
||||
@@ -199,7 +421,14 @@ The long form, with what was wrong before and how it was found, is in
|
||||
- Torrent enclosures through librqbit, seeding to a ratio or a time, with a stall timeout.
|
||||
- `ipx import` and `ipx export` for OPML, and systemd units in `contrib/`.
|
||||
|
||||
[unreleased]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.4.0...main
|
||||
[unreleased]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.6.0...main
|
||||
[0.6.0]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.5.5...v0.6.0
|
||||
[0.5.5]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.5.4...v0.5.5
|
||||
[0.5.4]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.5.3...v0.5.4
|
||||
[0.5.3]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.5.2...v0.5.3
|
||||
[0.5.2]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.5.1...v0.5.2
|
||||
[0.5.1]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.5.0...v0.5.1
|
||||
[0.5.0]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.4.0...v0.5.0
|
||||
[0.4.0]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.3.0...v0.4.0
|
||||
[0.3.0]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.2.0...v0.3.0
|
||||
[0.2.0]: https://git.sdf1.net/rays/ipodderx-rs/compare/v0.1.0...v0.2.0
|
||||
|
||||
35
CLAUDE.md
35
CLAUDE.md
@@ -17,6 +17,10 @@ Arcane project `content`: `/mnt/fast/arcane/projects/content/compose.yaml`. That
|
||||
| Database | `/mnt/user/ipodderx/state.db` | `/data/state.db` |
|
||||
| Downloads | `/mnt/user/ipodderx/downloads` | `/downloads` |
|
||||
| Web UI | `192.168.1.130:8099`, also `ipodderx.sdf1.net` via a Cloudflare tunnel | `0.0.0.0:8099` |
|
||||
| Sign-in via the tunnel | Cloudflare Access app `ipodderx`, with Authentik as its identity provider; see [docs/sso.md](docs/sso.md) | trusts `Cf-Access-Authenticated-User-Email` from `192.168.16.1`, the `content_default` gateway |
|
||||
|
||||
Work to do lives in the Gitea issues at https://git.sdf1.net/rays/ipodderx-rs/issues, not in a
|
||||
`TODO.md`. `/src/tea` is logged in: `/src/tea issues list --login git.sdf1.net --repo rays/ipodderx-rs`.
|
||||
|
||||
Deploying a change is: build and push the image, then pull it and recreate the container.
|
||||
|
||||
@@ -45,7 +49,10 @@ docker tag mirror.gcr.io/library/rust:1-slim-bookworm rust:1-slim-bookworm
|
||||
Run those again now and then, or the local copies go stale.
|
||||
|
||||
The healthcheck runs `ipx status` against the control socket, so `(healthy)` in `docker ps` means
|
||||
the worker is alive, not just the web port. The container restarts on its own after a reboot.
|
||||
the daemon answers there and can read its database, not just that the web port is up. The socket
|
||||
answers `status` itself instead of queuing it behind the worker's current job, so a long scan or
|
||||
download does not fail the check; it also means a worker stuck on one job would still pass. The
|
||||
container restarts on its own after a reboot.
|
||||
|
||||
Before the container, ipx ran by hand in code-server, with its files in `/config/.config/ipx/` and
|
||||
`/config/.local/share/ipx/`. Those are still there and the container does not read them. If you run
|
||||
@@ -99,9 +106,9 @@ Non-trivial logic leaves one runnable check behind. Pure functions (`merge_polic
|
||||
|
||||
* **`enclosures.url` is globally UNIQUE.** It is the dedupe key and the reason one file serves every
|
||||
subscriber. Two feeds publishing the same URL means only the first one scanned shows it.
|
||||
* **`entries.read`, `entries.flagged` and `entries.position` are dead columns.** Read state lives in
|
||||
`entry_state` per user. Two bugs have already come from queries still reading the old ones
|
||||
(retention, and the entry pruner) — grep before adding a third.
|
||||
* **Read state lives in `entry_state`, per user, and nowhere else.** `entries` had `read`, `flagged`
|
||||
and `position` columns from before accounts; two bugs came from queries still reading them
|
||||
(retention, and the entry pruner), and `migrate()` now drops them.
|
||||
* **The catalogue is config.toml; the subscriptions are in the database.** A feed exists once;
|
||||
`subscriptions(user_id, feed_id)` says who wants it and with what settings. OPML children are
|
||||
derived and never written to config.
|
||||
@@ -113,8 +120,13 @@ Non-trivial logic leaves one runnable check behind. Pure functions (`merge_polic
|
||||
watch the shutdown channel itself; the daemon ignored SIGTERM for exactly this reason.
|
||||
* Only one daemon per socket. Removing the socket file defeats the guard and you get two daemons
|
||||
fighting over the database, with the stale one still holding the port.
|
||||
* `/api/settings` answering `200` does **not** mean the worker is alive — it is a different task.
|
||||
Probe the control socket (`ipx status`) to check that.
|
||||
* `/api/settings` answering `200` does **not** mean the daemon is well — the web server is a
|
||||
different task. `ipx status` checks the control socket and the database; to see the worker
|
||||
getting through its jobs, watch for `scan complete` in the log.
|
||||
* **Every `ipx` command runs `migrate()` when it opens the database**, the healthcheck's
|
||||
`ipx status` included. A migration that rewrites a big table (`DROP COLUMN`) takes seconds on
|
||||
production, and a command run meanwhile fails with `migrating schema`. It changes nothing; wait
|
||||
for `daemon started` in the log. Copy `state.db` aside before deploying one.
|
||||
|
||||
## House style
|
||||
|
||||
@@ -142,3 +154,14 @@ Deliberate simplifications get a `ponytail:` comment naming the ceiling and the
|
||||
(documented in [docs/sso.md](docs/sso.md)).
|
||||
* A feed's `<description>` subtitle is dropped whenever `content:encoded` exists, which loses
|
||||
Substack-style subtitles.
|
||||
|
||||
<!-- rtk-instructions v2 -->
|
||||
# Command output
|
||||
|
||||
Command output here is condensed to save tokens, keeping every signal and
|
||||
dropping costly noise. Treat it as the complete result: run commands
|
||||
normally, and batch related commands into one call to avoid extra turns.
|
||||
Truncated results state their recovery path in their own output. Re-run a
|
||||
command as `rtk proxy <cmd>` only when its result is unusable: empty when
|
||||
output was clearly expected, contradicting its exit code, or garbled.
|
||||
<!-- /rtk-instructions -->
|
||||
34
Cargo.lock
generated
34
Cargo.lock
generated
@@ -436,17 +436,6 @@ dependencies = [
|
||||
"shlex",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "cfb"
|
||||
version = "0.14.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "a347dcabdae9c31b0825fd6a8bed285ec9c2acb89c47827126d52fa4f59cece3"
|
||||
dependencies = [
|
||||
"fnv",
|
||||
"uuid",
|
||||
"web-time",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "cfg-if"
|
||||
version = "1.0.4"
|
||||
@@ -867,15 +856,6 @@ dependencies = [
|
||||
"dirs-sys",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dirs"
|
||||
version = "7.0.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "8d57d423b3c82e89b9a24ca3091fee61f456a26edbd28d26c65906f4bc1dcd8f"
|
||||
dependencies = [
|
||||
"dirs-sys",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "dirs-sys"
|
||||
version = "0.5.0"
|
||||
@@ -1608,15 +1588,6 @@ dependencies = [
|
||||
"serde_core",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "infer"
|
||||
version = "0.22.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "f4200d433cbd5178df7797c9c2e75b348b728e39631cf14520d1e2fc424201f4"
|
||||
dependencies = [
|
||||
"cfb",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "intervaltree"
|
||||
version = "0.2.7"
|
||||
@@ -1634,7 +1605,7 @@ checksum = "791930b43c0d5973160d90a8f3894509f2b273430f5c5c73b668636d0287c5c0"
|
||||
|
||||
[[package]]
|
||||
name = "ipx"
|
||||
version = "0.4.0"
|
||||
version = "0.6.0"
|
||||
dependencies = [
|
||||
"ammonia",
|
||||
"anyhow",
|
||||
@@ -1643,9 +1614,7 @@ dependencies = [
|
||||
"axum",
|
||||
"chrono",
|
||||
"clap",
|
||||
"dirs",
|
||||
"futures-util",
|
||||
"infer",
|
||||
"librqbit",
|
||||
"opml",
|
||||
"percent-encoding",
|
||||
@@ -1656,7 +1625,6 @@ dependencies = [
|
||||
"serde",
|
||||
"serde_json",
|
||||
"tokio",
|
||||
"tokio-stream",
|
||||
"toml",
|
||||
"tower",
|
||||
"tower-http 0.7.1",
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "ipx"
|
||||
version = "0.4.0"
|
||||
version = "0.6.0"
|
||||
edition = "2024"
|
||||
|
||||
[dependencies]
|
||||
@@ -11,20 +11,17 @@ atom_syndication = "0.12.10"
|
||||
axum = "0.8.9"
|
||||
chrono = { version = "0.4.45", default-features = false, features = ["std", "clock"] }
|
||||
clap = { version = "4.6.6", features = ["derive"] }
|
||||
dirs = "7.0.0"
|
||||
futures-util = { version = "0.3.34", default-features = false, features = ["std"] }
|
||||
infer = "0.22.0"
|
||||
librqbit = { version = "9.0.1", default-features = false, features = ["rust-tls", "http-api-client"] }
|
||||
opml = "1.1.6"
|
||||
percent-encoding = "2.3.2"
|
||||
quick-xml = "0.42.0"
|
||||
quick-xml = { version = "0.42.0", features = ["escape-html"] }
|
||||
reqwest = { version = "0.13.5", default-features = false, features = ["rustls", "http2", "gzip", "stream", "json", "charset", "system-proxy"] }
|
||||
rss = "2.1.1"
|
||||
rusqlite = { version = "0.40.2", features = ["bundled"] }
|
||||
serde = { version = "1.0.229", features = ["derive"] }
|
||||
serde_json = "1.0.151"
|
||||
tokio = { version = "1.53.1", features = ["rt-multi-thread", "macros", "fs", "io-util", "net", "sync", "time", "signal"] }
|
||||
tokio-stream = { version = "0.1.19", features = ["sync"] }
|
||||
toml = "1.1.5"
|
||||
tower = { version = "0.5.3", features = ["util"] }
|
||||
tower-http = { version = "0.7.1", features = ["fs"] }
|
||||
|
||||
144
README.md
144
README.md
@@ -1,123 +1,83 @@
|
||||
# ipodderx-rs
|
||||
|
||||
A headless podcatcher: scans RSS/Atom feeds, downloads enclosures (HTTP and BitTorrent), files them
|
||||
into per-feed folders, and reaps old files to stay under a disk quota. Runs as a one-shot CLI or as
|
||||
a daemon with a web UI, serving any number of people from one copy of the data.
|
||||
A self-hosted podcatcher for a household. It checks your feeds, downloads the episodes, and serves
|
||||
a web UI modelled on the 2004 Mac app **iPodderX**, for any number of people sharing one copy of
|
||||
the files. One Rust binary, `ipx`, is both the daemon and the command line.
|
||||
|
||||
## Lineage
|
||||
It is a rewrite of [ipodderx-core](https://git.sdf1.net/rays/ipodderx-core), the Python engine
|
||||
behind iPodderX (2004-2008, Ray Slakinski & August Trometer).
|
||||
|
||||
A modern Rust rewrite of [ipodderx-core](https://git.sdf1.net/rays/ipodderx-core), the Python 2
|
||||
engine behind **iPodderX** (2004-2008, Ray Slakinski & August Trometer), open-sourced under the MIT
|
||||
License in 2010.
|
||||
## What it does
|
||||
|
||||
What carries over: the feed scan and TTL handling, GUID/URL dedupe, per-feed and per-date download
|
||||
folders, keyword filters, the explicit-content filter, torrent enclosures, and "SmartSpace" -- the
|
||||
oldest-first disk quota reaper.
|
||||
- **The web UI.** It has a toolbar, and a feed list that opens with Directory, Popular and All
|
||||
Subscriptions. Items sit in a sortable table with a Files pane, and there is a player bar. It
|
||||
comes in Dark, Light and Classic themes, and works on a phone.
|
||||
- **Several people, one copy.** Each person has their own subscriptions and their own read, pinned
|
||||
and playback state. There is one file on disk per episode, however many people want it. People
|
||||
sign in with a password or through a proxy (Cloudflare Zero Trust or Authentik), and admins
|
||||
manage accounts and settings.
|
||||
- **Scanning.** Feeds are checked on a schedule, globally or per feed, and a feed's own TTL is
|
||||
honoured. Keyword, explicit-content and media-type filters decide what is downloaded, with a cap
|
||||
on new downloads per scan.
|
||||
- **Downloads.** Files come over HTTP or BitTorrent and are filed into a folder per feed.
|
||||
Retention deletes the oldest files to stay under a disk quota or an age limit, and never touches
|
||||
an item someone has pinned.
|
||||
- **OPML.** You can import and export your own subscriptions. You can also subscribe to an OPML
|
||||
URL, which keeps a whole list in step as a folder.
|
||||
|
||||
What does not: iTunes and iPhoto export via AppleScript, text-to-speech enclosures, the Windows
|
||||
WMP/COM paths, XML plists and Python pickles for state, the `directory.iPodderX.com` survey ping,
|
||||
3DES-encrypted preferences, and the `printMSG` stdout protocol -- replaced by a JSON-lines socket.
|
||||
## Run it
|
||||
|
||||
## Quick start
|
||||
With Docker:
|
||||
|
||||
```sh
|
||||
docker build -t ipodderx .
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
`docker-compose.yml` is set up for the author's own server. Point its `image` and its three volumes
|
||||
(`/config`, `/data` and `/downloads`) at yours first. The UI is on port 8099. BitTorrent uses 6881
|
||||
over TCP and UDP. Files are written as `PUID`/`PGID`, 99:100 by default.
|
||||
|
||||
From source:
|
||||
|
||||
```sh
|
||||
cargo build --release
|
||||
install -m755 target/release/ipx ~/.cargo/bin/
|
||||
|
||||
ipx add https://atp.fm/rss # subscribe
|
||||
ipx fetch # scan and download
|
||||
ipx daemon # scheduler, control socket and web UI
|
||||
./target/release/ipx daemon
|
||||
```
|
||||
|
||||
On first start with `[web] enabled = true` the daemon mints a token, writes it to config.toml and
|
||||
prints the URL to open. A database with no accounts starts with **admin / ipodderx** at `/login` --
|
||||
change it with `echo -n '<password>' | ipx user passwd admin`.
|
||||
The first start creates **admin / ipodderx**. Sign in at `/login`, then change it:
|
||||
|
||||
```sh
|
||||
echo -n 'a good password' | ipx user passwd admin
|
||||
```
|
||||
|
||||
The UI is plain HTTP, so put TLS in front of it if it is reachable from outside your network.
|
||||
|
||||
## Documentation
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| [docs/configuration.md](docs/configuration.md) | Every config key, paths, environment variables |
|
||||
| [docs/configuration.md](docs/configuration.md) | Every config key, path and environment variable |
|
||||
| [docs/cli.md](docs/cli.md) | Every command, including `ipx user` |
|
||||
| [docs/users.md](docs/users.md) | Accounts, and what several people share |
|
||||
| [docs/sso.md](docs/sso.md) | Cloudflare Zero Trust or Authentik in front |
|
||||
| [docs/architecture.md](docs/architecture.md) | How it works: modules, schema, socket, HTTP API |
|
||||
| [docs/sso.md](docs/sso.md) | Signing in through Cloudflare Zero Trust or Authentik |
|
||||
| [docs/architecture.md](docs/architecture.md) | How it works: modules, schema, control socket, HTTP API |
|
||||
| [CHANGELOG.md](CHANGELOG.md) | What changed, by release |
|
||||
| [docs/history.md](docs/history.md) | How it was built: the long form, with what was wrong and why |
|
||||
| [CLAUDE.md](CLAUDE.md) | Notes for anyone (or anything) working on the code |
|
||||
|
||||
## The web UI
|
||||
|
||||
`ipx daemon` serves it in the same process, so it reads SQLite and the event bus directly.
|
||||
|
||||
Feeds down the side; the selected feed's items across the top; the selected item's text and its
|
||||
enclosures below, which is where you play, download or delete them. The divider drags and its
|
||||
position is remembered. Playback serves Range requests, so seeking works. An OPML subscription is a
|
||||
collapsible folder whose page lists the feeds inside it.
|
||||
|
||||
An item may carry several enclosures; all of them appear below, and anything that is not audio or
|
||||
video gets a View link rather than a player -- the publisher's copy until it is downloaded, the
|
||||
local one after. Opening an item marks it read. Show notes are untrusted feed HTML, sanitized with
|
||||
`ammonia` server-side before they reach the page.
|
||||
|
||||
The **Log** button shows the running daemon live in four tabs: *Daemon I/O* is the control protocol
|
||||
itself, every command in and event out; *Scans* is feed and download activity; *HTTP* is web
|
||||
requests; *All* is everything, with level and text filters and a copy button. It reads a ring buffer
|
||||
held in the process, not a file, so it works the same under Docker.
|
||||
|
||||
It is plain HTTP. On a LAN bind the token and everything else cross the network in the clear, and a
|
||||
feed URL can itself carry a credential. Put TLS in front of it if that matters.
|
||||
|
||||
## OPML
|
||||
|
||||
**Importing and exporting** a file copies subscriptions in or out once: `ipx import subs.opml`,
|
||||
`ipx export subs.opml`, or Settings → Subscriptions in the UI.
|
||||
|
||||
**Subscribing to an OPML URL** is a live subscription, as iPodderX had. Add the OPML's URL like any
|
||||
other feed; every scan re-reads it and keeps your list in step. The feeds inside are not written to
|
||||
config.toml -- the OPML is the source of truth, so they are re-derived each scan and held in the
|
||||
database. They show as a folder, download into one nested folder, and inherit the subscription's
|
||||
settings until you change one, which gives it its own entry.
|
||||
|
||||
When a feed drops out of the OPML upstream, it is unsubscribed and removed -- unless it has
|
||||
downloads, in which case it is kept and flagged in the UI as no longer listed. A downloaded file is
|
||||
never left behind with nothing explaining where it came from.
|
||||
|
||||
## Docker
|
||||
|
||||
```sh
|
||||
docker buildx build --tag 192.168.1.130:5000/ipodderx:latest . --push
|
||||
docker compose pull ipodderx && docker compose up -d ipodderx
|
||||
docker compose logs -f ipodderx # the first start prints the default admin password
|
||||
```
|
||||
|
||||
`docker-compose.yml` runs the image from the registry above rather than building it, so build and
|
||||
push first; change the tag in both places to use another registry. It mounts `/config` (config.toml),
|
||||
`/data` (state.db) and `/downloads` from this install's host paths, which you will want to change for
|
||||
yours. It publishes 8099 for the
|
||||
UI and 6881 (TCP **and** UDP -- DHT needs the UDP side), and sets `PUID`/`PGID` to `99:100` so files
|
||||
land owned the way Unraid shares expect. The healthcheck runs `ipx status` through the control
|
||||
socket, so it catches a daemon that is alive but wedged rather than merely one that has died.
|
||||
|
||||
## Running it as a service
|
||||
|
||||
`contrib/` has a systemd user unit for the daemon, and a timer plus one-shot service if you would
|
||||
rather run periodic scans with no daemon -- in which case there is no socket for a UI to attach to.
|
||||
| [docs/history.md](docs/history.md) | How it was built, with what was wrong and why |
|
||||
| [CLAUDE.md](CLAUDE.md) | Notes for working on the code, including how production is deployed |
|
||||
|
||||
## Tests
|
||||
|
||||
```sh
|
||||
cargo test # the engine: parsing, filters, retention, schedules, SQL, per-user state
|
||||
node tests/page-smoke.js # the page script loads without throwing
|
||||
npx playwright test # a real browser against a real daemon
|
||||
npx playwright test # a real browser against a real daemon on fixture feeds
|
||||
```
|
||||
|
||||
`npm install` gets the test runner; the browser comes from
|
||||
`npx playwright install --with-deps chromium` (in `install.sh`).
|
||||
`npm install` gets the test runner, and `npx playwright install --with-deps chromium` gets the
|
||||
browser.
|
||||
|
||||
## License
|
||||
|
||||
MIT. See [LICENSE](LICENSE).
|
||||
|
||||
The icons are [Font Awesome Free](https://fontawesome.com) 7.3.1 by @fontawesome, under
|
||||
[CC BY 4.0](https://fontawesome.com/license/free), embedded as SVG in `web/index.html`.
|
||||
MIT, see [LICENSE](LICENSE). The icons are [Font Awesome Free](https://fontawesome.com) 7.3.1 by
|
||||
@fontawesome, under [CC BY 4.0](https://fontawesome.com/license/free), embedded as SVG.
|
||||
|
||||
@@ -1,7 +0,0 @@
|
||||
[Unit]
|
||||
Description=ipx feed scan (one shot)
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=%h/.cargo/bin/ipx fetch
|
||||
Environment=IPX_LOG=ipx=info
|
||||
@@ -1,18 +0,0 @@
|
||||
# User unit: install to ~/.config/systemd/user/ipx.service, then
|
||||
# systemctl --user enable --now ipx
|
||||
# The socket lands in $XDG_RUNTIME_DIR/ipx.sock by default, so a UI running as the
|
||||
# same user can attach without extra configuration.
|
||||
[Unit]
|
||||
Description=ipx podcatcher
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
ExecStart=%h/.cargo/bin/ipx daemon
|
||||
Restart=on-failure
|
||||
RestartSec=30
|
||||
Environment=IPX_LOG=ipx=info
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -1,16 +0,0 @@
|
||||
# Alternative to the daemon: a periodic one-shot scan, closer to how the original
|
||||
# iPodderX agent was driven. Use this OR ipx.service, not both -- with no daemon
|
||||
# running there is no socket, so a UI cannot attach.
|
||||
#
|
||||
# Install ipx-scan.service and ipx.timer to ~/.config/systemd/user/, then
|
||||
# systemctl --user enable --now ipx.timer
|
||||
[Unit]
|
||||
Description=Periodic ipx feed scan
|
||||
|
||||
[Timer]
|
||||
OnBootSec=5min
|
||||
OnUnitActiveSec=1h
|
||||
Persistent=true
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -28,12 +28,15 @@ The page is compiled in, so **editing `web/index.html` needs a rebuild**.
|
||||
1. Skip the feed unless `last_checked + max(schedule, ttl)` has passed (`--force` ignores this).
|
||||
2. Conditional GET with the stored `ETag` / `Last-Modified`. `304` ends it there.
|
||||
3. Sniff the body: RSS, then Atom, then OPML. An OPML is a live subscription — its feeds are
|
||||
re-derived into the database each scan, never written to config.toml.
|
||||
re-derived into the database each scan, never written to config.toml. A Patreon creator link
|
||||
(a token, no `show=`) with more than one show is treated the same way, before any fetch: its
|
||||
shows come from Patreon's web API and each becomes a derived feed.
|
||||
4. Record entries. A changed title or description flips the item back to unread.
|
||||
5. Record enclosures. `enclosures.url` is `UNIQUE`, which is the dedupe key and subsumes the
|
||||
original's `history.dat` pickle: a reaped file keeps its row so it is never fetched twice.
|
||||
6. Apply the merged policy (see [users.md](users.md)) and mark anything rejected as `skipped` with
|
||||
a reason.
|
||||
a reason. What a filter skipped is judged again every scan, so a change of settings brings it
|
||||
back. A feed in a group takes your settings on the group for anything you have not set on it.
|
||||
7. Download what is still pending, newest first, up to the per-scan cap. A `.torrent` body goes to
|
||||
the torrent path whatever its advertised type; an HTML body is a failed download — a login wall
|
||||
or an error page — and is deleted.
|
||||
@@ -47,21 +50,23 @@ entries feed_id, guid, title, link, published, description, first_seen,
|
||||
image, duration, episode, season PK (feed_id, guid)
|
||||
enclosures id, feed_id, guid, url UNIQUE, mime, length, path, state,
|
||||
bytes_done, downloaded_at, last_error
|
||||
users id, name, pass_hash, is_admin, created
|
||||
sessions token, user_id, created, seen
|
||||
users id, name, pass_hash, is_admin, created, last_login
|
||||
sessions token, user_id, seen
|
||||
subscriptions user_id, feed_id, keywords, auto_download, allow_explicit,
|
||||
max_new_per_check, created PK (user_id, feed_id)
|
||||
max_new_per_check PK (user_id, feed_id)
|
||||
entry_state user_id, feed_id, guid, read, flagged, position
|
||||
PK (user_id, feed_id, guid)
|
||||
```
|
||||
|
||||
`entries` still has `read`, `flagged` and `position` columns from before accounts existed. They are
|
||||
**dead** — the migration copied them into `entry_state` and nothing reads them now. Anything found
|
||||
querying them is a bug; two were.
|
||||
Read state is `entry_state` alone. `entries` had `read`, `flagged` and `position` columns from
|
||||
before accounts; two bugs came from queries still reading them, and `migrate()` drops them from an
|
||||
older database.
|
||||
|
||||
Schema changes: add the table or column to `SCHEMA`, and for a column also to the list in
|
||||
`migrate()`, which does `PRAGMA table_info` then `ALTER TABLE ADD COLUMN`. `Db::memory()` runs the
|
||||
same path as `Db::open`, so a migration-only column cannot pass tests while missing in production.
|
||||
Schema changes: add the table or column to `SCHEMA`. `CREATE TABLE IF NOT EXISTS` leaves a table
|
||||
that already exists alone, so a new column on one also goes in `migrate()`'s `wanted` list, and a
|
||||
retired one in its `retired` list; both are checked with `PRAGMA table_info`. Columns from before
|
||||
0.3.0, the oldest version an upgrade may start from, need no entry. `Db::memory()` runs the same
|
||||
path as `Db::open`, so a migration cannot pass the tests while missing in production.
|
||||
|
||||
## Control socket
|
||||
|
||||
@@ -82,7 +87,9 @@ printf '{"cmd":"fetch","force":true}\n' | socat - UNIX-CONNECT:$XDG_RUNTIME_DIR/
|
||||
**Events** — `feed_start`, `feed_skip`, `feed_done`, `feed_error`, `progress`, `download_done`,
|
||||
`download_error`, `torrent_deferred`, `reaped`, `reap_done`, `scan_done`, `status`, `error`.
|
||||
`scan_done`, `reap_done` and `status` are terminal: a client that asked for work stops reading
|
||||
there.
|
||||
there. Commands run one at a time, in the order they arrive, except `status`: the socket answers it
|
||||
straight away, so the Docker healthcheck is never left waiting behind a scan or a download, and
|
||||
answers only the client that asked, since `status` would end any other client's session.
|
||||
|
||||
Progress carries the enclosure id, without which a UI cannot tell one download from another and
|
||||
ends up animating every pending row. It is throttled to whole percents. The stream is a broadcast,
|
||||
@@ -108,11 +115,11 @@ else a `401`.
|
||||
| `GET /api/entries` | the same, across every feed you subscribe to (All Subscriptions) |
|
||||
| `POST /api/feeds/{id}/read-all`, `POST /api/feeds/{id}/download-latest` | |
|
||||
| `POST /api/read-all` | everything read in every feed you subscribe to (All Subscriptions) |
|
||||
| `POST /api/entries/{feed}/{guid}/flags`, `…/position` | your read, starred, position |
|
||||
| `POST /api/entries/{feed}/{guid}/flags`, `…/position` | your read, kept, position |
|
||||
| `POST /api/enclosures/{id}/download`, `DELETE /api/enclosures/{id}` | `?force=true` overrides the shared-file warning |
|
||||
| `POST /api/fetch` | |
|
||||
| `GET /api/opml`, `POST /api/opml` | export your subscriptions; subscribe to every feed in an OPML |
|
||||
| `GET /api/popular`, `GET /api/directory`, `POST /api/popular/{id}` | the ten most subscribed feeds, and every listable feed A to Z, with everyone counted (id, title, art, count, whether it is yours; never a URL, never a private feed); subscribe by id |
|
||||
| `GET /api/popular`, `GET /api/directory`, `POST /api/popular/{id}` | the ten most subscribed feeds, and every listable feed A to Z, with an OPML's feeds in place of the OPML and everyone counted (id, title, art, count, whether it is yours, the feed's iTunes category, whether it carries audio or video; never a URL, never a private feed); subscribe by id |
|
||||
| `GET /api/settings`, `PATCH /api/settings` | admin-only to write |
|
||||
| `GET /api/users`, `POST /api/users`, `PATCH /api/users/{id}`, `DELETE /api/users/{id}` | admin-only; the only admin cannot be demoted or removed |
|
||||
| `GET /api/events` | SSE, the same broadcast the socket carries |
|
||||
|
||||
@@ -60,7 +60,7 @@ ipx reap # actually delete
|
||||
```
|
||||
|
||||
Files are deleted to get back under `max_total_gb`, oldest first, and items past `max_age_days`
|
||||
with no file are pruned from the database. **Starred by anyone keeps a file**, and one only counts
|
||||
with no file are pruned from the database. **An item anyone kept keeps its file**, and one only counts
|
||||
as read when everyone subscribed has read it. The enclosure row survives as `reaped`, which is what
|
||||
stops the next scan fetching it again.
|
||||
|
||||
|
||||
@@ -32,10 +32,10 @@ media_types = ["audio", "video"]
|
||||
be polled *less* often, and a per-feed `schedule` overrides both. Admin-only from the UI.
|
||||
* **`organize`** — `feed` files downloads under the feed's folder; `date` under `YYYY-MM-DD`.
|
||||
* **`max_total_gb`** — the reaper deletes to get back under this, oldest first, keeping a 50 MB
|
||||
pad. Starred items are never deleted, and a file only counts as read once every subscriber has
|
||||
pad. Kept items are never deleted, and a file only counts as read once every subscriber has
|
||||
read it. `0` disables it entirely.
|
||||
* **`max_age_days`** — items older than this with no file on disk are pruned from the database.
|
||||
Starred ones stay. `0` disables it.
|
||||
Kept ones stay. `0` disables it.
|
||||
* **`max_new_per_check`** — the cap that stops a new subscription pulling a whole back catalogue.
|
||||
`0` means unlimited, which is rarely what you want: subscribing to an OPML of 80 feeds with no cap
|
||||
fetched 216 files and 22 GB in one scan.
|
||||
@@ -43,8 +43,6 @@ media_types = ["audio", "video"]
|
||||
can be fetched by hand; blog feeds put each article's header image in an `<enclosure>`, and
|
||||
without this the disk fills with artwork. Empty takes everything.
|
||||
|
||||
`interval_mins` from older configs is still read, and `schedule` supersedes it.
|
||||
|
||||
## `[torrent]`
|
||||
|
||||
```toml
|
||||
@@ -69,6 +67,7 @@ token = "" # generated and saved on first run
|
||||
trusted_header = "" # e.g. "Cf-Access-Authenticated-User-Email"
|
||||
trusted_proxies = ["127.0.0.1", "::1"]
|
||||
auto_create_users = true
|
||||
sign_out_url = "" # e.g. "/cdn-cgi/access/logout"
|
||||
session_days = 30
|
||||
```
|
||||
|
||||
@@ -79,6 +78,9 @@ session_days = 30
|
||||
* **`trusted_proxies`** — addresses allowed to assert that header, and the entire security boundary
|
||||
for it. Name the proxy, never a subnet.
|
||||
* **`auto_create_users`** — create an account the first time the proxy vouches for a new name.
|
||||
* **`sign_out_url`** — where Sign out sends someone the proxy signed in: the proxy's own sign-out,
|
||||
`/cdn-cgi/access/logout` behind Cloudflare Access. Empty sends them to the sign-in page, where
|
||||
the proxy signs them straight back in.
|
||||
* **`session_days`** — sign a session out after this long without a request.
|
||||
|
||||
It is plain HTTP. On a LAN bind everything crosses the network in the clear — and a feed URL can
|
||||
@@ -95,6 +97,7 @@ url = "https://atp.fm/rss"
|
||||
folder = "Accidental Tech Podcast" # default: the feed title
|
||||
schedule = "every 6h" # overrides [general] for this feed
|
||||
media_types = ["audio"] # overrides [general] for this feed
|
||||
category = "Technology" # the Directory's, if the feed names none
|
||||
username = "ray" # HTTP basic auth
|
||||
password_env = "IPX_ATP_PASS" # preferred over a literal `password`
|
||||
```
|
||||
|
||||
286
docs/history.md
286
docs/history.md
@@ -6,6 +6,292 @@ reasoning lives. New write-ups go at the top.
|
||||
|
||||
See [README.md](../README.md) for what the thing is.
|
||||
|
||||
## 2026-09-15 — Currently Listening, empty for anyone who opens what they play
|
||||
|
||||
Issue #14: an episode 32 minutes into 41 was missing from Currently Listening, which said nothing
|
||||
was in progress. The filter was "position past five seconds and unread", on the reasoning that
|
||||
`markPlayed` marks an episode read at 90%. But opening an item marks it read too, and you open an
|
||||
episode to play it. In production every one of the twelve episodes with a saved position was read,
|
||||
so the list was empty for everyone. The browser test for it had set the episode unread before
|
||||
checking, to cope with an earlier test having opened it, and so tested around exactly this.
|
||||
|
||||
Finished now means the saved position is 90% of the episode's length or more, the same line
|
||||
`markPlayed` draws, and read plays no part. Three of those twelve episodes had no length in their
|
||||
feed, and with no length there is no telling finished from started, so they would have stayed
|
||||
listed for good. The player now sends the length it measured with each saved position, and that
|
||||
fills in a missing one, never replacing a length the feed gave.
|
||||
|
||||
## 2026-09-15 — davewiner's 922 rows, retired at last
|
||||
|
||||
davewiner's OPML subscription left config.toml before `retire_group` existed, so nothing ever
|
||||
retired the 922 feeds derived from it. The backstop in `subscriptions()` kept them from being
|
||||
scanned, but the rows stayed, and 56 of the 75 errors stored in production were theirs: stale, and
|
||||
never going to change.
|
||||
|
||||
Retiring them the obvious way would have done damage. Eleven of the 922 -- xkcd, The Verge,
|
||||
TechCrunch, Hacker News and others -- had since been given config entries of their own and were
|
||||
scanned from there, but their rows still said `managed = 1`. `managed_feeds()` returned them, so
|
||||
`retire_group("davewiner")` would have dropped the ten with no files as derived feeds, deleting
|
||||
the entries people were reading. `unmanage` exists for exactly this and had never been applied.
|
||||
|
||||
`retire_group` now unmanages a feed that has its own config entry instead of dropping it, which
|
||||
also covers removing any OPML or Patreon subscription from the page, and the daemon retires every
|
||||
group whose parent is gone from config when it starts. In production that is davewiner alone: 792
|
||||
rows with nothing downloaded forgotten, 119 with files kept as orphaned, 11 unmanaged. One of the
|
||||
792, `ars-technica-all-content-2`, has a subscriber but neither a config entry nor a file; the
|
||||
backstop had already hidden it, so nobody could see it to lose it.
|
||||
|
||||
## 2026-09-14 — Settings, for everyone with an account
|
||||
|
||||
A user reported that Settings disappeared shortly after they signed in: it showed for a moment,
|
||||
then was gone. `#prefs` sat inside the same `.tgroup` as `#logs`, and `api('/api/me')` hid the
|
||||
whole group -- `$('#admintools').hidden=true` -- the moment it learned the account was not an
|
||||
admin. Nothing wrong with that check timing; it was hiding the wrong thing.
|
||||
|
||||
The Settings modal is not actually all-or-nothing. `GET /api/settings`, and Export and Import
|
||||
OPML, carry no admin check server-side -- `export_opml` and `import_opml` work from a user's own
|
||||
subscriptions, and the schedule/quota page is read-only information, not a control. Only the
|
||||
`PATCH` that changes those settings, and the Users screen behind it, return 403 for anyone but an
|
||||
admin. The comment above the old hide -- "scanning, quotas, accounts and the log are the
|
||||
operator's business" -- was wrong about quotas and half wrong about accounts: reading them is
|
||||
everyone's; changing them is the operator's.
|
||||
|
||||
`prefsModal()` now branches on `S.me.admin` the way the per-feed settings modal already does for
|
||||
its URL field: a non-admin gets the schedule and quota as text, Subscriptions (Export/Import)
|
||||
in full, and no Users section or Save button. Only `#logs` stays hidden, since the log names every
|
||||
account and every failed sign-in. The browser test for a second account asserted the old
|
||||
behaviour outright (`#prefs` hidden, not an admin) rather than what the server actually allows;
|
||||
fixing the UI meant fixing the test's premise too, not just the assertion.
|
||||
|
||||
## 2026-09-12 — Healthy while busy
|
||||
|
||||
After a deploy the container sat at "starting" for a minute, and Docker's health log showed two
|
||||
`ipx status` probes exceeding their 5-second timeout. The daemon's own log explained it. The first
|
||||
scan after the start fetched 23 feeds, from 14:10:41 to 14:11:35, and both probes' `status`
|
||||
commands waited in the job queue behind it; they were answered together at 14:11:35, straight after
|
||||
`scan_done`. The worker runs one job at a time and `status` was one of its jobs, so any scan or
|
||||
download longer than about a minute and a half, three 30-second probes, would have had Docker call
|
||||
a working daemon unhealthy.
|
||||
|
||||
The socket now answers `status` itself, from two short queries, and only real work goes through the
|
||||
queue. The trade is that healthy now means the daemon answers on its socket and can read its
|
||||
database; a worker stuck on one job would still pass. Asking a daemon that downloads hour-long
|
||||
podcasts to be idle within five seconds was never a fair test of whether it was alive. A test holds
|
||||
the queue full and checks `status` still comes back.
|
||||
|
||||
The first version broadcast the answer, as the queued one had been. Timing `status` during a forced
|
||||
scan in production showed the catch: `status` is a terminal event, so the `ipx fetch` watching that
|
||||
scan stopped reading at the first probe and printed the status line as its last, while the scan
|
||||
carried on. When `status` waited behind the scan it could never arrive first, so this had never
|
||||
shown. The answer now goes only to the client that asked, and the test checks that another client
|
||||
hears nothing.
|
||||
|
||||
## 2026-09-12 — Signing in through Authentik, for real
|
||||
|
||||
Ray could not get Authentik's sign-in to reach ipx, following `docs/sso.md`, which had been written
|
||||
without ever being tried. Looking at the Cloudflare account through its API showed that side was
|
||||
already complete. Authentik is Zero Trust's OpenID Connect identity provider; the Access application
|
||||
`ipodderx` allows only it and a list of five addresses; the tunnel `rays-unraid` routes
|
||||
`ipodderx.sdf1.net` to `192.168.1.130:8099`; DNS is a proxied CNAME to the tunnel. Access's log
|
||||
showed `rays@sdf1.net` signing in through it. Nothing on Cloudflare was changed, so no other site
|
||||
was touched.
|
||||
|
||||
The gaps were all at ipx's end: `trusted_header` was empty, `trusted_proxies` held only loopback,
|
||||
and the account was called `rays` while the header carries `rays@sdf1.net`.
|
||||
|
||||
Finding the address to trust took the most time. The page said `127.0.0.1`, but `cloudflared` runs in
|
||||
its own container and reaches ipx through the host's published port. ipx logs no peer addresses, so
|
||||
the address was read from `/proc/net/tcp` inside the ipx container: `192.168.16.1`, the gateway of
|
||||
`content_default`, where Docker's masquerade puts traffic crossing from another bridge. A request
|
||||
from Tower's own shell arrived as `192.168.1.130` instead, and a throwaway `busybox` on the default
|
||||
bridge as `192.168.16.1`: the first was refused with the header, the second believed. LAN machines
|
||||
keep their own addresses, since Docker forwards published ports with iptables (the userland proxy
|
||||
only handles loopback).
|
||||
|
||||
Every change, in order, with how to undo it:
|
||||
|
||||
1. **Code**, commit `586d2c0`: `ipx user rename`, deployed. Revert the commit and redeploy to
|
||||
remove it; nothing depends on it once used.
|
||||
2. **Account**: `docker exec iPodderX ipx user rename rays rays@sdf1.net`. Same id, so its feeds,
|
||||
read state, password and admin rights stayed. Undo: `docker exec iPodderX ipx user rename
|
||||
rays@sdf1.net rays`. Signing in at `/login` now takes the new name.
|
||||
3. **Config**, `/mnt/fast/appdata/ipodderx/config.toml`, `[web]`: `trusted_header` from `""` to
|
||||
`"Cf-Access-Authenticated-User-Email"`, and `"192.168.16.1"` added to `trusted_proxies`. The
|
||||
file as it was is `config.toml.2026-09-12-sso.bak` beside it. Undo: copy the backup back and
|
||||
`docker compose -f /mnt/fast/arcane/projects/content/compose.yaml restart ipodderx`.
|
||||
4. **Cloudflare, Docker networks and other containers**: unchanged. The `busybox` test container
|
||||
was removed when it exited, and its image afterwards.
|
||||
5. **Authentik**, later the same day, because ipodderx had no tile in its library while Outline
|
||||
did: a bookmark application `ipodderx` (pk `5854a98e-816a-4c4f-9f27-63e69dc29d1d`), made
|
||||
through the API with a token of Ray's. No provider and no policy bindings, like Outline's, the
|
||||
iPodderX icon, and a link to `https://ipodderx.sdf1.net`. It changes nothing about who can sign
|
||||
in. Undo: delete it under Applications → Applications, or
|
||||
`DELETE /api/v3/core/applications/ipodderx/`.
|
||||
6. **Signing out**, later again. Sign out landed on ipx's password page while Access still vouched
|
||||
for Ray, so it signed nothing out, and the page looked like the wrong login. Cloudflare's
|
||||
`/cdn-cgi/access/logout` ends the Access session for every Access application at once (there is
|
||||
no per-application sign-out, and it takes no redirect), and Authentik's end-session only ends
|
||||
one application's session unless single logout is set up there. Ray chose Access's sign-out. New
|
||||
`[web] sign_out_url`, set to `/cdn-cgi/access/logout` in production (the file as it was is
|
||||
`config.toml.2026-09-12-signout.bak`), and `/login` now sends anyone the proxy vouches for on to
|
||||
`/`. Undo: take the key out and restart; the code does nothing without it.
|
||||
|
||||
What the address trusts is any container on Tower that connects through the host's port, not only
|
||||
`cloudflared`. Verifying Cloudflare's signed `Cf-Access-Jwt-Assertion` would remove that, and is
|
||||
the upgrade if it matters.
|
||||
|
||||
## 2026-09-12 — Trimming the state database
|
||||
|
||||
An audit of the database layer, with a read-only copy of production to check it against. The
|
||||
data was already clean: no tables or indexes left from older versions, 47 free pages after the
|
||||
column drops earlier the same day, and one stray `entry_state` row. The code had five things:
|
||||
|
||||
- `migrate()` still added eight columns to any table missing them. All eight shipped in 0.2.0 and
|
||||
upgrades now start from 0.3.0 at the oldest, so the list and its loop went; the `retired` drop
|
||||
list stays, since a database coming from 0.4.0 still has the old read columns.
|
||||
- `created` on `users`, `subscriptions` and `sessions` was written by every insert and read by
|
||||
nothing. They joined `retired`. The old-database test now builds all three tables, foreign keys
|
||||
included, since `DROP COLUMN` on a table that references another was the part worth proving.
|
||||
- `Db::subscribed_feed_ids` had no callers, though its doc said the scanner walked it.
|
||||
`Db::subscriber_count` had one caller asking whether it was above zero, which
|
||||
`subscriber_counts().contains_key` answers. `Managed.orphaned` was selected and never read.
|
||||
- `users.created` came back the same afternoon, with `last_login` beside it. Nothing read it, but
|
||||
when an account was made and when it last signed in is what you want to know when tidying
|
||||
accounts, and it cannot be recovered later. Both existing accounts got their creation times back
|
||||
from the backup taken before the drop, and a last sign-in from their newest session in it.
|
||||
`last_login` is kept to the hour, because the proxy vouches for every request and that would
|
||||
otherwise be a write each time.
|
||||
|
||||
## 2026-09-12 — Cutting what had outlived its reason
|
||||
|
||||
A whole-repo audit for over-engineering listed twelve things to cut, and all of them went.
|
||||
|
||||
- **Upgrades from before accounts.** `migrate_opml_children` moved OPML feeds that old versions
|
||||
wrote into `config.toml` out to the database, and ran at every daemon start to do nothing after
|
||||
the first. Production ran it in 0.3.0; anything older has to pass through 0.4.0.
|
||||
- **Half of the adoption, and not the other half.** The audit called `adopt_existing_library` a
|
||||
one-time migration and it was cut whole. It did two jobs: copy the old read state into
|
||||
`entry_state`, which was dead, and subscribe the first admin to the whole catalogue while nobody
|
||||
subscribed to anything, which is how a fresh install's first account gets `config.toml`'s feeds.
|
||||
The browser suite caught it at once, signing in to an empty sidebar; `cargo test` had no idea.
|
||||
The second job is back as `adopt_catalogue`, with a unit test of its own.
|
||||
- **The dead `entries` columns.** `read`, `flagged` and `position` moved to `entry_state` with
|
||||
accounts. The adoption's copy was their last reader, but `record_entry` still wrote them, and
|
||||
still reset `read` when a title changed, which nothing looked at. Two bugs came from queries
|
||||
reading them. `migrate()` now drops them from an existing database (SQLite has had `DROP COLUMN`
|
||||
since 3.35), and a test builds an old table to prove it. On production each drop rewrote the
|
||||
66 MB `entries` table, about four seconds apiece, so the first start took thirteen. An
|
||||
`ipx status` run in that window failed with `migrating schema`: every `ipx` command migrates when
|
||||
it opens the database, and it collided with the daemon doing the same. A failed `ALTER TABLE`
|
||||
changes nothing, and the database had been copied to `backup/` first anyway.
|
||||
- **`interval_mins`**, which `schedule` replaced. An old config that still has the key loads; the
|
||||
key is ignored, and the config test carries it to keep that true.
|
||||
- **Three dependencies.** `infer` was only asked whether a file is a torrent, and the check after
|
||||
it already looked for `d8:announce`, which is what `infer` looks for. `dirs` was three lookups of
|
||||
`XDG_CONFIG_HOME`, `XDG_DATA_HOME` and `HOME`. `tokio-stream` wrapped the broadcast receiver for
|
||||
the event stream; `futures_util::stream::unfold` does the same, lagging clients included.
|
||||
- **Two token generators.** The web token came from a copy of the session-token code, with a
|
||||
clock fallback on top. It uses `auth::new_session_token` now, and is 64 characters.
|
||||
- **The icon inlined four times**, 23 KB of base64 each, into both pages. It is `/icon.png` now,
|
||||
outside the sign-in wall with `/login`, since the sign-in page shows it.
|
||||
- Also: the `contrib/` systemd units from before the container, `Db::entries` and
|
||||
`Db::count_entries` that only the tests called, three `logbuf` visitors that repeated the trait's
|
||||
defaults, and unused state, a helper and dead CSS in the page.
|
||||
|
||||
## 2026-09-11 — A design pass on the web UI
|
||||
|
||||
A review against screenshots of every view in all three themes found that Dark and Light read as a
|
||||
generic dark dashboard: one pale blue did every job, most labels were bold, and nothing led. The
|
||||
list it produced, in `TODO.md`, was worked through in one go. What was worth knowing:
|
||||
|
||||
- **Amber means new.** Badges, unread dots and download bars take the icon's EQ amber; blue is left
|
||||
for the primary action and links. Light's amber was `#b06f10`, which gives white text 4.1:1,
|
||||
short of AA for 11 px bold. It is `#9a5f0a` now, 5.2:1.
|
||||
- **EQ bars mark what is playing.** Three `<i>` bars stand at 60, 100 and 40 % and animate only
|
||||
while `body.playing` is set. The first version left the animation on but paused, expecting each
|
||||
bar to hold a different frame. The frames it held were within a pixel of each other, and on
|
||||
screen the bars read as three dots. Under reduced motion one rule drops every animation and
|
||||
transition, which leaves the bars standing.
|
||||
- **The focus ring was clipped.** `.tgroup` and `#topbar` both set `overflow:hidden`, so a ring
|
||||
drawn outside a toolbar button was cut off. Rings inside clipping parents are inset instead.
|
||||
- **The feed list could not be used from the keyboard at all.** Rows were `<div>`s and the triangle
|
||||
a `<span>`, so Tab went from the feed filter to Sign out. Rows now take focus, the triangle is a
|
||||
`<button aria-expanded>`, and `renderFeeds` puts focus back on the same row after redrawing,
|
||||
since every live update replaces every row. The list's own key handler stops Space and the
|
||||
arrows from reaching the player's shortcuts on the document.
|
||||
- **The triangle hangs in the margin.** Every row used to reserve an 18 px slot for it, pushing a
|
||||
hundred feeds 28 px right of the places above for the sake of two folders. It is now absolutely
|
||||
placed in the row's left padding, the full height of the row, so a near miss no longer opens
|
||||
the folder's page.
|
||||
- **A selected tile vanished** because the initials tile and the selected row were both `--raise`.
|
||||
Tiles now mix their tint into `--bg`, which no row uses.
|
||||
- **The Files pane hides itself** with `#split:has(>#files[hidden])`, which collapses its column.
|
||||
The phone layout already hides the pane, and a zero-width extra track there is harmless.
|
||||
- **Flagged became Kept** in the tab, and in the server's refusal to delete a file someone else
|
||||
kept. The filter value and the column stay `flagged`; renaming those buys nothing.
|
||||
|
||||
## 2026-09-11 — A Patreon creator is a list of shows
|
||||
|
||||
Ray asked whether ipx could sync with Patreon. Not in full. The documented API (v2, the
|
||||
`identity.memberships` scope) lists the creators you back and whether each has a feed (`has_rss`),
|
||||
but no resource carries the `auth` token that makes a feed URL work. That token only comes from the
|
||||
creator's page. It is also one per membership, not one per account: techpod's differs from Glass
|
||||
Cannon's, so no single token finds everything you back.
|
||||
|
||||
What does work is one creator at a time, which is what Ray wanted for Glass Cannon and its 33 shows:
|
||||
|
||||
- `patreon.com/rss?auth=<token>`, with no creator named, returns that token's creator. Its self link,
|
||||
about 660 bytes in, gives the campaign by number (`/rss/369921`). Patreon ignores `Range` here, so
|
||||
ipx reads the stream until the number appears and hangs up, instead of taking all 2.8 MB.
|
||||
- A show's `show=` number is a Patreon collection. Asked anonymously, the collection listing
|
||||
(`/api/collection?filter[campaign_id]=`) and a post's `collections` both hide patron-only ones: you
|
||||
get "FAQ". `/api/campaigns/<id>?include=shows` lists every show, anonymously, in one response.
|
||||
- Every spelling works: `rss/glasscannon?auth=…&show=N`, `rss/369921?…` and `rss?auth=…&show=N` all
|
||||
serve the same 131 items. Enclosure URLs are the same in the creator feed and the show feed, and
|
||||
stable between fetches.
|
||||
|
||||
That last point shaped the design. `enclosures.url` is unique, so whichever feed is scanned first owns
|
||||
the file. The first cut only asked a creator for its shows while it had no entries of its own, so that a
|
||||
creator already read as a plain feed, holding every show's episodes, would never be split into shows
|
||||
that came up empty. Within the hour that was the wrong call: Glass Cannon had gone into production on
|
||||
the build before this one, been read as one feed of 2,385 items, and the rule kept it that way. Finding
|
||||
anything in that heap was the problem Ray wanted solved.
|
||||
|
||||
So a creator with more than one show is always a group, run through the same sync as an OPML
|
||||
(`sync_group`, split out of `sync_opml`). When it becomes one, its items are cleared and each show
|
||||
takes over the enclosures the creator holds as the show lists them (`Db::adopt`), downloaded files and
|
||||
everyone's read state included. One show leaves it a plain feed, which is what techpod already was. If
|
||||
the shows cannot be listed, a creator already split fails the scan rather than being read as one heap;
|
||||
one that never was is read as one feed until they can be. An answer without a `shows` list is an error,
|
||||
not "no shows".
|
||||
|
||||
**Filter verdicts follow the settings.** Ray also reported that turning on Allow explicit and
|
||||
rescanning brought nothing back. An item was judged once, when first seen, and `skipped` was final. The
|
||||
2026-09-10 entry below saw it coming ("worth a `ipx retry <feed>` command if this bites"). It bit:
|
||||
2,166 Glass Cannon items and all 88 of Shadowdark's sat at `skipped: explicit` with the setting on.
|
||||
Every scan now runs the filters again over what they skipped (not over `torrents disabled`, which is not
|
||||
a filter's call) and requeues what they now let through. Only that direction: a queued item is never
|
||||
pulled back, because Download latest and a manual download both work by queueing.
|
||||
|
||||
**Two gaps beside it.** Add feed had no explicit box, so every new feed's first scan skipped all its
|
||||
explicit items; it has one now, stored on your subscription like the feed dialog's. And a feed inside a
|
||||
group ignored your settings on the group, though the group's dialog said they were inherited: settings
|
||||
live on each person's subscription, and nothing read the group's. `Db::subscribers` now fills what you
|
||||
have not set on the feed from your subscription to the group, and the feed list shows the same.
|
||||
|
||||
**A name that was already used.** Replaying the split on a copy of the production database left one
|
||||
show with a Supercast episode in it. "Glass Cannon Live! Ascension | Pathfinder 2E" slugs to
|
||||
`glass-cannon-live-ascension-pathfinder-2`, the id of a Supercast feed of the same show that had been
|
||||
removed. Removing a feed keeps its rows on purpose, so that re-adding it does not fetch the back
|
||||
catalogue again, but choosing a new id only checked config.toml and derived feeds. The Patreon show took
|
||||
the old id and everything still filed under it. An id is now also taken when the database has a feed by
|
||||
that id at a different URL; the same URL may still have it back, which is the re-add case.
|
||||
|
||||
Shows already added by hand are matched by token and show number, not by exact URL (`same_feed`), so a
|
||||
bare token does not add Get in the Trunk and Shadowdark a second time under another spelling.
|
||||
|
||||
The show listing is Patreon's own undocumented web API. If it changes, only finding new shows stops.
|
||||
|
||||
## 2026-09-11 — One meaning per icon, sortable columns, and one player
|
||||
|
||||
Ray asked for a pass over the whole UI: consistent icons, and buttons placed next to what they act
|
||||
|
||||
309
docs/sso.md
309
docs/sso.md
@@ -1,9 +1,8 @@
|
||||
# Signing in through Cloudflare Zero Trust or Authentik
|
||||
# Signing in through Cloudflare Access and Authentik
|
||||
|
||||
ipx can take the signed-in identity from whatever sits in front of it, instead of asking for a
|
||||
password itself. Both products below do the same thing in the end: they authenticate the person and
|
||||
pass the result to the origin in a **header**. ipx reads that header, finds (or creates) the
|
||||
matching account, and gets on with it.
|
||||
password itself. The proxy authenticates the person and passes the result to ipx in a **header**;
|
||||
ipx reads it, finds (or creates) the matching account, and gets on with it.
|
||||
|
||||
Read [How this is secured](#how-this-is-secured) before exposing anything. The short version: a
|
||||
header is worth exactly as much as the hop that set it, so ipx only believes one from an address you
|
||||
@@ -11,195 +10,173 @@ list.
|
||||
|
||||
---
|
||||
|
||||
## The ipx side (both setups)
|
||||
## How ipodderx.sdf1.net does it
|
||||
|
||||
Checked end to end on 2026-09-12. An earlier version of this page had never been tried against a
|
||||
real setup and pointed at the wrong address.
|
||||
|
||||
```
|
||||
browser ─► Cloudflare Access, app "ipodderx" ─── sign in ───► Authentik (OpenID Connect)
|
||||
─► tunnel "rays-unraid" (the cloudflared container on Tower)
|
||||
─► http://192.168.1.130:8099 ─► ipx
|
||||
```
|
||||
|
||||
Authentik is not in the request path. It is the identity provider Cloudflare Access asks. Access
|
||||
then adds `Cf-Access-Authenticated-User-Email`, the email address Authentik gave it, to every
|
||||
request it forwards through the tunnel, and ipx signs that person in.
|
||||
|
||||
| Piece | Where | Setting |
|
||||
|---|---|---|
|
||||
| Identity provider | Zero Trust → Settings → Authentication | `Authentik`, OpenID Connect; scopes `openid email profile` |
|
||||
| Access application | Zero Trust → Access → Applications → `ipodderx` | Domain `ipodderx.sdf1.net`; identity providers: Authentik only, with instant auth; session 730h; policy *Require Login* allows a list of email addresses |
|
||||
| Tunnel route | Zero Trust → Networks → Tunnels → `rays-unraid` → Public hostnames | `ipodderx.sdf1.net` → HTTP `192.168.1.130:8099` |
|
||||
| DNS | `sdf1.net` | `ipodderx` CNAME to the tunnel, proxied |
|
||||
| ipx | `/mnt/fast/appdata/ipodderx/config.toml`, `[web]` | below |
|
||||
|
||||
```toml
|
||||
[web]
|
||||
enabled = true
|
||||
bind = "0.0.0.0:8099"
|
||||
token = "…" # keep it: it is the admin, used by the healthcheck
|
||||
|
||||
# The header your proxy sets. Empty (the default) disables this whole path.
|
||||
trusted_header = "Cf-Access-Authenticated-User-Email" # Authentik: "X-authentik-username"
|
||||
|
||||
# Addresses allowed to assert that header -- the proxy, and nothing else.
|
||||
trusted_proxies = ["127.0.0.1", "::1"]
|
||||
|
||||
# Create an account the first time the proxy vouches for a name ipx has not seen.
|
||||
bind = "0.0.0.0:8099"
|
||||
trusted_header = "Cf-Access-Authenticated-User-Email"
|
||||
trusted_proxies = ["127.0.0.1", "::1", "192.168.16.1"]
|
||||
auto_create_users = true
|
||||
|
||||
sign_out_url = "/cdn-cgi/access/logout"
|
||||
session_days = 30
|
||||
```
|
||||
|
||||
Restart the daemon after editing. Accounts made this way have **no password**: they can only ever
|
||||
arrive through the proxy. `ipx user list` marks them `proxy only`.
|
||||
Restart ipx after editing it: `docker compose -f /mnt/fast/arcane/projects/content/compose.yaml
|
||||
restart ipodderx`.
|
||||
|
||||
The first account created is an admin. Every later one is an ordinary user, and an ordinary user
|
||||
cannot change global settings, a feed's URL or folder, or how often feeds are scanned: the API
|
||||
refuses those with a `403`, not just the UI. Everything else about a feed (which items they want,
|
||||
whether to fetch them, how many at a time) is theirs alone; see [users.md](users.md).
|
||||
### What was missing
|
||||
|
||||
Somebody arriving through the proxy for the first time starts with **no feeds**, because
|
||||
subscriptions are per person. Adding a feed someone else already reads costs no second fetch and no
|
||||
second copy on disk.
|
||||
Cloudflare and Authentik were already right. Three things on the ipx side were not:
|
||||
|
||||
Promote someone with:
|
||||
1. **`trusted_header` was empty**, which switches the whole proxy path off. ipx ignored the header
|
||||
and asked for a password.
|
||||
2. **`trusted_proxies` listed only `127.0.0.1`.** The tunnel's requests do not come from there;
|
||||
see the next section.
|
||||
3. **The account had the wrong name.** It was made by hand as `rays`, but the header carries
|
||||
`rays@sdf1.net`. With `auto_create_users` on, the first visit would have made a second, empty
|
||||
account. `ipx user rename rays rays@sdf1.net` fixed that without losing anything.
|
||||
|
||||
### The address to trust, and why it is 192.168.16.1
|
||||
|
||||
`cloudflared` runs in its own container and reaches ipx through the host's published port. Docker
|
||||
(iptables firewall backend) masquerades traffic between its bridge networks, so the tunnel's
|
||||
requests arrive from the **gateway of ipx's own network**, `content_default`:
|
||||
|
||||
```sh
|
||||
ipx user list
|
||||
echo -n 'a good password' | ipx user passwd <name> # optional: also lets them sign in directly
|
||||
docker network inspect content_default -f '{{range .IPAM.Config}}{{.Gateway}}{{end}}'
|
||||
```
|
||||
|
||||
Local sign-in at `/login` keeps working alongside all of this, which is how you get in from the LAN
|
||||
when the tunnel is down. So does the shared `[web] token`, which signs in as the admin: that is
|
||||
what the Docker healthcheck uses, and the way back in if you lock yourself out. A brand new database
|
||||
starts with **admin / ipodderx** — change it.
|
||||
That was measured, not assumed. ipx does not log where a request came from, so the addresses were
|
||||
read from the kernel's connection table inside the container while the site was open. (`/proc/net/tcp`
|
||||
lists them in hex.)
|
||||
|
||||
If the `content` project's network is ever recreated, its gateway can change. Check it again, and
|
||||
update `trusted_proxies` to match.
|
||||
|
||||
### Names
|
||||
|
||||
The username is the email address, lower-cased: `rays@sdf1.net`. To sign in at `/login` with a
|
||||
password from the LAN, use that name too.
|
||||
|
||||
To let someone else in, add their address to the Access policy; they need an Authentik account with
|
||||
that email. With `auto_create_users = true` they get an ipx account on their first visit, as an
|
||||
ordinary user with no feeds. An account made before the proxy can be given the name the proxy will
|
||||
send:
|
||||
|
||||
```sh
|
||||
docker exec iPodderX ipx user rename <old name> <email address>
|
||||
```
|
||||
|
||||
### Signing out
|
||||
|
||||
**Sign out** sends someone the proxy signed in to `sign_out_url`, here Cloudflare's
|
||||
`/cdn-cgi/access/logout`. That ends your Access session for **every** Access application,
|
||||
`code.sdf1.net` included: Cloudflare has no way to end just one, and its sign-out page does not send
|
||||
you anywhere afterwards. The next visit goes back through Authentik, which lets you straight in if
|
||||
you are still signed in there. Signing out of Authentik itself is Authentik's own sign-out.
|
||||
|
||||
ipx never shows its password page to someone the proxy vouches for: `/login` sends them on to their
|
||||
feeds.
|
||||
|
||||
### The tile in Authentik's library
|
||||
|
||||
Authentik's library lists Authentik's own applications, and ipodderx signs in through the one
|
||||
called `Cloudflare Access`, so ipodderx needs a bookmark of its own to show up there. It is
|
||||
Applications → Applications → `ipodderx`: no provider, launch URL `https://ipodderx.sdf1.net`, and
|
||||
the iPodderX icon. Like Outline's, it has no policy bindings, so everyone in Authentik sees the
|
||||
tile. Who actually gets in is still up to the Access policy.
|
||||
|
||||
### Check it
|
||||
|
||||
```sh
|
||||
# From Tower itself: not a trusted address, so the header is ignored.
|
||||
curl -s -H 'Accept: application/json' -H 'Cf-Access-Authenticated-User-Email: rays@sdf1.net' \
|
||||
http://192.168.1.130:8099/api/me # -> sign in
|
||||
|
||||
# From a container on a Docker bridge, as cloudflared is: believed.
|
||||
docker run --rm --network bridge mirror.gcr.io/library/busybox wget -qO- \
|
||||
--header 'Accept: application/json' --header 'Cf-Access-Authenticated-User-Email: rays@sdf1.net' \
|
||||
http://192.168.1.130:8099/api/me # -> {"admin":true,"name":"rays@sdf1.net"}
|
||||
```
|
||||
|
||||
Then open `https://ipodderx.sdf1.net` in a private window. Authentik should ask who you are, and
|
||||
ipx should show `rays@sdf1.net` in the sidebar footer without asking for a password.
|
||||
|
||||
---
|
||||
|
||||
## Cloudflare Zero Trust
|
||||
## The ipx settings
|
||||
|
||||
This is what runs `ipodderx.sdf1.net`: a `cloudflared` tunnel to the origin, with an Access
|
||||
application in front of it. Cloudflare authenticates the visitor and adds
|
||||
`Cf-Access-Authenticated-User-Email` to every request it forwards.
|
||||
|
||||
### 1. The tunnel
|
||||
|
||||
In **Zero Trust → Networks → Tunnels**, either use the existing tunnel or create one, then add a
|
||||
public hostname:
|
||||
|
||||
| Field | Value |
|
||||
| Key | What it does |
|
||||
|---|---|
|
||||
| Subdomain / domain | `ipodderx` / `sdf1.net` |
|
||||
| Type | HTTP |
|
||||
| URL | `localhost:8099` (or the LAN address of the box) |
|
||||
| `trusted_header` | The header the proxy sets. Empty, the default, turns the proxy path off. |
|
||||
| `trusted_proxies` | The addresses allowed to set it. Nothing else is believed. |
|
||||
| `auto_create_users` | Make an account the first time the proxy vouches for a name ipx has not seen. |
|
||||
| `sign_out_url` | Where Sign out sends someone the proxy signed in: the proxy's own sign-out. Empty sends them to the sign-in page, where the proxy signs them straight back in. |
|
||||
| `session_days` | How long a password sign-in lasts without use. |
|
||||
|
||||
Use `localhost` when `cloudflared` runs on the same machine as ipx — that keeps the origin request
|
||||
coming from `127.0.0.1`, which is already in `trusted_proxies`. If `cloudflared` runs elsewhere (its
|
||||
own container, another host), put **its** address in `trusted_proxies` instead, and make sure
|
||||
nothing else can reach port 8099.
|
||||
The first account ever created is an admin. Every later one is an ordinary user, who cannot change
|
||||
global settings, a feed's URL or folder, or how often feeds are scanned: the API refuses those with
|
||||
a `403`, not just the UI. Everything else about a feed is theirs alone; see [users.md](users.md).
|
||||
|
||||
### 2. The Access application
|
||||
|
||||
**Zero Trust → Access → Applications → Add an application → Self-hosted**:
|
||||
|
||||
- Application domain: `ipodderx.sdf1.net`
|
||||
- Session duration: whatever suits; ipx keeps its own 30-day session on top.
|
||||
- Add a policy — *Allow*, with a rule such as `Emails` → your address, or `Emails ending in` →
|
||||
your domain. Anyone this policy admits gets an ipx account when `auto_create_users` is on, so keep
|
||||
the policy as narrow as the people you actually want reading your feeds.
|
||||
|
||||
### 3. Point ipx at the header
|
||||
|
||||
```toml
|
||||
trusted_header = "Cf-Access-Authenticated-User-Email"
|
||||
trusted_proxies = ["127.0.0.1", "::1"]
|
||||
```
|
||||
|
||||
The username becomes the email address, lower-cased (`ray@example.com`). That is what shows in the
|
||||
sidebar and what `ipx user list` prints.
|
||||
|
||||
### 4. Check it
|
||||
|
||||
```sh
|
||||
# From the box itself: no header, no session -> the sign-in page.
|
||||
curl -s -o /dev/null -w '%{http_code} %{redirect_url}\n' -H 'Accept: text/html' http://127.0.0.1:8099/
|
||||
|
||||
# Pretending to be the tunnel (only works because 127.0.0.1 is trusted):
|
||||
curl -s -H 'Cf-Access-Authenticated-User-Email: you@example.com' http://127.0.0.1:8099/api/me
|
||||
```
|
||||
|
||||
Then load `https://ipodderx.sdf1.net` in a browser: Cloudflare should ask who you are, and ipx
|
||||
should show your address in the sidebar footer without ever asking for a password.
|
||||
Local sign-in at `/login` keeps working alongside the proxy, which is how you get in from the LAN
|
||||
when the tunnel is down. So does the shared `[web] token`, which signs in as the admin and is the
|
||||
way back in if you lock yourself out. A brand new database starts with **admin / ipodderx**;
|
||||
change it.
|
||||
|
||||
---
|
||||
|
||||
## Authentik
|
||||
## Authentik in the request path instead
|
||||
|
||||
Authentik does this with a **Proxy Provider** plus an **outpost**, which sits in the request path and
|
||||
adds `X-authentik-username` (also `X-authentik-email`, `X-authentik-name`, `X-authentik-groups`).
|
||||
Not what ipodderx.sdf1.net uses, and **not verified**. Authentik can also sit in front of ipx
|
||||
itself, with a **Proxy Provider** and an **outpost** that adds `X-authentik-username`:
|
||||
|
||||
### 1. Provider
|
||||
|
||||
**Applications → Providers → Create → Proxy Provider**:
|
||||
|
||||
- Name: `ipx`
|
||||
- Authorization flow: your usual (`default-provider-authorization-implicit-consent`)
|
||||
- Mode: **Forward auth (single application)** if an existing reverse proxy fronts ipx, or
|
||||
**Proxy** to let the outpost talk to ipx directly.
|
||||
- External host: `https://ipodderx.example.net`
|
||||
- Internal host (Proxy mode): `http://<ip of the ipx box>:8099`
|
||||
|
||||
### 2. Application and outpost
|
||||
|
||||
**Applications → Create**, bind it to that provider, and give it a policy so only the people you
|
||||
mean are let through. Then add the provider to an outpost (**Applications → Outposts**, the embedded
|
||||
one is fine).
|
||||
|
||||
### 3. Forward auth, if you use nginx/SWAG in front
|
||||
|
||||
In the server block for ipx:
|
||||
|
||||
```nginx
|
||||
location /outpost.goauthentik.io {
|
||||
proxy_pass http://authentik-server:9000/outpost.goauthentik.io;
|
||||
proxy_set_header Host $host;
|
||||
proxy_set_header X-Original-URL $scheme://$http_host$request_uri;
|
||||
add_header Set-Cookie $auth_cookie;
|
||||
auth_request_set $auth_cookie $upstream_http_set_cookie;
|
||||
}
|
||||
|
||||
location / {
|
||||
auth_request /outpost.goauthentik.io/auth/nginx;
|
||||
error_page 401 = @goauthentik_proxy_signin;
|
||||
auth_request_set $auth_cookie $upstream_http_set_cookie;
|
||||
add_header Set-Cookie $auth_cookie;
|
||||
|
||||
# This is the line that matters to ipx.
|
||||
auth_request_set $authentik_username $upstream_http_x_authentik_username;
|
||||
proxy_set_header X-authentik-username $authentik_username;
|
||||
|
||||
proxy_pass http://ipx:8099;
|
||||
}
|
||||
```
|
||||
|
||||
### 4. Point ipx at the header
|
||||
|
||||
```toml
|
||||
trusted_header = "X-authentik-username"
|
||||
trusted_proxies = ["172.18.0.5"] # the outpost or nginx container, NOT a whole subnet
|
||||
```
|
||||
|
||||
Usernames arrive as Authentik knows them (`ray`), lower-cased.
|
||||
- Applications → Providers → Create → Proxy Provider; mode **Proxy** (the outpost talks to ipx) or
|
||||
**Forward auth** (an existing reverse proxy asks the outpost).
|
||||
- Applications → Create, bound to that provider, with a policy; add the provider to an outpost.
|
||||
- In ipx: `trusted_header = "X-authentik-username"`, and the outpost's or reverse proxy's address
|
||||
in `trusted_proxies`. Measure that address as above rather than guessing it.
|
||||
|
||||
---
|
||||
|
||||
## How this is secured
|
||||
|
||||
**The header is only believed from `trusted_proxies`.** Every other source is ignored, and the
|
||||
request falls through to a session cookie or the shared token. This is the whole security boundary,
|
||||
so:
|
||||
request falls through to a session cookie or the shared token. That is the whole security boundary.
|
||||
|
||||
- List the **proxy's own address**, not a range. `["127.0.0.1"]` when the tunnel runs beside ipx;
|
||||
the container's IP when it does not.
|
||||
- Never list a LAN subnet. Anyone on your network could then send
|
||||
`Cf-Access-Authenticated-User-Email: admin@…` and be your admin.
|
||||
- Make sure the origin port is not reachable *around* the proxy by anyone you would not admit
|
||||
through it. If it is, bind ipx to `127.0.0.1` and let only the proxy reach it.
|
||||
With the tunnel reaching ipx through the host's port, `192.168.16.1` means **any container on Tower
|
||||
that connects to `192.168.1.130:8099`**, not only `cloudflared`. Machines on the LAN, and Tower
|
||||
itself, arrive under their own addresses and cannot set the header; the checks above show both
|
||||
sides. Never list a LAN address or range: anyone there could then send
|
||||
`Cf-Access-Authenticated-User-Email: rays@sdf1.net` and be you.
|
||||
|
||||
Verify the refusal, don't assume it — set `trusted_proxies = ["10.9.9.9"]` briefly and confirm a
|
||||
header from your machine gets a `401`:
|
||||
**What ipx does not do:** it does not verify Cloudflare's signed `Cf-Access-Jwt-Assertion`. It
|
||||
trusts the hop. Verifying the signature would make the containers on Tower irrelevant to the
|
||||
boundary, and is the upgrade if that ever matters.
|
||||
|
||||
```sh
|
||||
curl -s -o /dev/null -w '%{http_code}\n' \
|
||||
-H 'Cf-Access-Authenticated-User-Email: someone@example.com' http://127.0.0.1:8099/api/me
|
||||
```
|
||||
|
||||
**What ipx does not do:** it does not verify Cloudflare's `Cf-Access-Jwt-Assertion` signature or
|
||||
Authentik's session. It trusts the hop. That is a deliberate trade — it keeps the configuration to
|
||||
three lines — and it is sound exactly as long as the point above holds.
|
||||
|
||||
**Turning it off:** clear `trusted_header`. Existing proxy-only accounts stay, but nobody can sign
|
||||
**Turning it off:** clear `trusted_header` and restart. Proxy-made accounts stay, but nobody can sign
|
||||
in with them until they are given a password (`ipx user passwd <name>`).
|
||||
|
||||
---
|
||||
@@ -207,22 +184,20 @@ in with them until they are given a password (`ipx user passwd <name>`).
|
||||
## Everyday administration
|
||||
|
||||
```sh
|
||||
ipx user list # who exists, and how each one signs in
|
||||
ipx user list # who exists, how each signs in, and when
|
||||
echo -n 'secret123' | ipx user add sam # local account, password on stdin
|
||||
ipx user add sam --no-password # proxy-only account, created ahead of time
|
||||
ipx user add sam@example.com --no-password # proxy-only account, made ahead of time
|
||||
ipx user rename sam sam@example.com # give an account the name the proxy sends
|
||||
echo -n 'newsecret' | ipx user passwd sam # change a password
|
||||
ipx user rm sam # remove the account
|
||||
```
|
||||
|
||||
Set `auto_create_users = false` once everyone who should have an account has one. After that the
|
||||
proxy vouching for an unknown name is logged and refused, rather than quietly making an account.
|
||||
Pre-create people instead with `ipx user add <name> --no-password`, using exactly the name the
|
||||
header will carry (Cloudflare sends the email address, lower-cased).
|
||||
In the container, put `docker exec iPodderX` in front, and `docker exec -i iPodderX` for the ones
|
||||
that read a password.
|
||||
|
||||
Scanning intervals, the disk quota, retention, the download folder and a feed's URL are
|
||||
**admin-only**: the Settings button is hidden for everyone else, and the API refuses the change even
|
||||
if the request is made by hand. Everyone controls their own keywords, auto-download, explicit
|
||||
setting and per-scan cap, along with their own read state and which feeds they see.
|
||||
Set `auto_create_users = false` once everyone who should have an account has one. After that the
|
||||
proxy vouching for an unknown name is logged and refused. Make people ahead of time instead, with
|
||||
the exact name the header will carry.
|
||||
|
||||
See also [users.md](users.md) for what several people share, [configuration.md](configuration.md)
|
||||
for every `[web]` key, and [cli.md](cli.md) for the `ipx user` commands.
|
||||
|
||||
@@ -8,7 +8,7 @@ fetch, one parse and one file.
|
||||
|
||||
| Yours alone | The same for everyone |
|
||||
|---|---|
|
||||
| Read, starred, playback position | The feed's URL |
|
||||
| Read, pinned, playback position | The feed's URL |
|
||||
| Which feeds you see at all | Its download folder |
|
||||
| Keywords, auto-download, explicit, per-scan cap | When it is scanned |
|
||||
| | The file on disk |
|
||||
@@ -34,10 +34,10 @@ is shared.
|
||||
|
||||
Deleting a file deletes everyone's copy. A feed with other subscribers labels the button **Delete
|
||||
for everyone** and names them in the confirmation, and the server has the last word: if anyone else
|
||||
has starred the item or not played it yet, `DELETE /api/enclosures/{id}` answers `409` with the
|
||||
has pinned the item or not played it yet, `DELETE /api/enclosures/{id}` answers `409` with the
|
||||
reason, and only `?force=true` goes through.
|
||||
|
||||
Retention follows the same rule: starred by anyone keeps a file, and it counts as read only once
|
||||
Retention follows the same rule: an item anyone pinned keeps its file, and it counts as read only once
|
||||
every subscriber has read it.
|
||||
|
||||
## Signing in
|
||||
@@ -72,13 +72,16 @@ re-subscribing does not pull the back catalogue again.
|
||||
|
||||
**Popular** and **Directory** sit at the top of the feed list, above your own feeds. Popular, also
|
||||
shown in the Add feed dialog, lists the ten feeds with the most subscribers on this server, you
|
||||
included. Directory lists every one of them A to Z. Your own feeds are marked Subscribed.
|
||||
It shows a title, artwork and a count, never a URL or who reads it. Feeds from an
|
||||
OPML subscription are left out, since they come with the OPML. So is anything that looks private: a
|
||||
login configured for the feed, credentials in its URL, or a key such as `auth=` or `token=` in the
|
||||
query, or a feed from a paid-feed service such as Patreon or Supercast, which put the key in the
|
||||
path. Those are someone's paid subscriptions, and listing them would let anyone here read what they
|
||||
pay for.
|
||||
included. Directory shows every one of them A to Z as a grid of cover art. Above it, a filter
|
||||
picks Podcasts (anything with audio or video) or Blogs (the rest), and chips pick the category
|
||||
each show gives itself in iTunes; the two combine. Your own feeds are marked Subscribed.
|
||||
It shows a title, artwork and a count, never a URL or who reads it. An OPML subscription is listed
|
||||
as the feeds inside it, one by one, and never the OPML itself, so you can take just the shows you
|
||||
want. Anything that looks private is left out: a login configured for the feed, credentials in its URL,
|
||||
or a key such as `auth=` or `token=` in the query, or a feed from a paid-feed service such as
|
||||
Patreon or Supercast, which put the key in the path, and any feed inside an OPML that looks private
|
||||
itself. Those are someone's paid subscriptions, and listing them would let anyone here read what
|
||||
they pay for.
|
||||
|
||||
An admin can do the same from **Settings → Manage users…**: add someone (with a password, or none
|
||||
for someone the proxy signs in), tick or untick Admin, or remove an account. Removing one takes its
|
||||
|
||||
11
skills-lock.json
Normal file
11
skills-lock.json
Normal file
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"version": 1,
|
||||
"skills": {
|
||||
"security-audit": {
|
||||
"source": "cloudflare/security-audit-skill",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/security-audit/SKILL.md",
|
||||
"computedHash": "98b97aad2873b3b9a8e064007c25e7827b493ba27fda8bfbdb41dab586767973"
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -26,9 +26,6 @@ pub struct General {
|
||||
/// How often to re-check feeds: "every 30m", "every 4h", "90" (minutes), "1d".
|
||||
/// A feed's own `schedule` overrides this.
|
||||
pub schedule: String,
|
||||
/// Superseded by `schedule`. Still read so existing configs keep working.
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub interval_mins: Option<u64>,
|
||||
pub organize: Organize,
|
||||
/// 0 = unlimited.
|
||||
pub max_total_gb: f64,
|
||||
@@ -83,6 +80,10 @@ pub struct Web {
|
||||
pub trusted_proxies: Vec<String>,
|
||||
/// Create an account the first time the proxy vouches for a name it has not seen.
|
||||
pub auto_create_users: bool,
|
||||
/// Where Sign out sends someone the proxy signed in. Signing out of ipx alone cannot stick
|
||||
/// while the proxy still vouches for them, so this is the proxy's own sign-out:
|
||||
/// `/cdn-cgi/access/logout` behind Cloudflare Access. Empty sends them to /login.
|
||||
pub sign_out_url: String,
|
||||
/// Sign a session out after this long without a request.
|
||||
pub session_days: i64,
|
||||
}
|
||||
@@ -96,6 +97,7 @@ impl Default for Web {
|
||||
trusted_header: String::new(),
|
||||
trusted_proxies: vec!["127.0.0.1".into(), "::1".into()],
|
||||
auto_create_users: true,
|
||||
sign_out_url: String::new(),
|
||||
session_days: 30,
|
||||
}
|
||||
}
|
||||
@@ -113,6 +115,10 @@ pub struct Feed {
|
||||
/// Download folder name; defaults to the sanitized feed title.
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub folder: Option<String>,
|
||||
/// The Directory's category for a feed that names none of its own, as most blogs do not.
|
||||
/// The feed's own iTunes category wins where there is one.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub category: Option<String>,
|
||||
/// Every whitespace-separated word of a keyword must appear in the
|
||||
/// url/title/description/categories for an enclosure to be taken.
|
||||
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
||||
@@ -153,7 +159,6 @@ impl Default for General {
|
||||
download_dir: home().join("Podcasts"),
|
||||
socket: default_socket(),
|
||||
schedule: "every 60m".into(),
|
||||
interval_mins: None,
|
||||
organize: Organize::Feed,
|
||||
max_total_gb: 0.0,
|
||||
max_age_days: 0,
|
||||
@@ -176,8 +181,8 @@ impl Default for Torrent {
|
||||
}
|
||||
|
||||
impl General {
|
||||
/// Minutes between checks. Falls back to the legacy `interval_mins`, then to an hour.
|
||||
/// A malformed value warns rather than stopping the daemon.
|
||||
/// Minutes between checks, or an hour when `schedule` is empty or unreadable. A malformed
|
||||
/// value warns rather than stopping the daemon.
|
||||
pub fn interval(&self) -> u64 {
|
||||
if let Some(n) = parse_interval(&self.schedule) {
|
||||
return n;
|
||||
@@ -185,7 +190,7 @@ impl General {
|
||||
if !self.schedule.trim().is_empty() {
|
||||
tracing::warn!(schedule = %self.schedule, "unrecognised schedule; using the default");
|
||||
}
|
||||
self.interval_mins.filter(|n| *n > 0).unwrap_or(60)
|
||||
60
|
||||
}
|
||||
}
|
||||
|
||||
@@ -278,9 +283,7 @@ pub fn config_path() -> PathBuf {
|
||||
if let Ok(p) = std::env::var("IPX_CONFIG") {
|
||||
return PathBuf::from(p);
|
||||
}
|
||||
dirs::config_dir()
|
||||
.unwrap_or_else(|| home().join(".config"))
|
||||
.join("ipx/config.toml")
|
||||
xdg("XDG_CONFIG_HOME", ".config").join("ipx/config.toml")
|
||||
}
|
||||
|
||||
/// `$IPX_DATA_DIR`, else `$XDG_DATA_HOME/ipx`.
|
||||
@@ -288,9 +291,7 @@ pub fn data_dir() -> PathBuf {
|
||||
if let Ok(p) = std::env::var("IPX_DATA_DIR") {
|
||||
return PathBuf::from(p);
|
||||
}
|
||||
dirs::data_dir()
|
||||
.unwrap_or_else(|| home().join(".local/share"))
|
||||
.join("ipx")
|
||||
xdg("XDG_DATA_HOME", ".local/share").join("ipx")
|
||||
}
|
||||
|
||||
fn default_socket() -> PathBuf {
|
||||
@@ -346,8 +347,16 @@ pub fn unique_slug(text: &str, taken: &BTreeMap<String, Feed>) -> String {
|
||||
(2..).map(|n| format!("{base}-{n}")).find(|s| !taken.contains_key(s)).unwrap()
|
||||
}
|
||||
|
||||
/// `$var`, or `~/fallback` when it is unset or empty, as the XDG base directory spec says.
|
||||
fn xdg(var: &str, fallback: &str) -> PathBuf {
|
||||
std::env::var_os(var)
|
||||
.filter(|v| !v.is_empty())
|
||||
.map(PathBuf::from)
|
||||
.unwrap_or_else(|| home().join(fallback))
|
||||
}
|
||||
|
||||
fn home() -> PathBuf {
|
||||
dirs::home_dir().unwrap_or_else(|| PathBuf::from("."))
|
||||
std::env::var_os("HOME").map(PathBuf::from).unwrap_or_else(|| PathBuf::from("."))
|
||||
}
|
||||
|
||||
fn expand_tilde(p: &Path) -> PathBuf {
|
||||
@@ -367,6 +376,8 @@ mod tests {
|
||||
r#"
|
||||
[general]
|
||||
download_dir = "/tmp/pods"
|
||||
# A key older versions read. An old config that still has it has to load.
|
||||
interval_mins = 45
|
||||
|
||||
[feeds.example]
|
||||
url = "https://example.com/feed.xml"
|
||||
@@ -403,22 +414,17 @@ mod tests {
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn interval_falls_back_through_legacy_then_default() {
|
||||
fn interval_falls_back_to_an_hour() {
|
||||
let mut g = General::default();
|
||||
assert_eq!(g.interval(), 60, "the default schedule");
|
||||
|
||||
g.schedule = "every 15m".into();
|
||||
assert_eq!(g.interval(), 15);
|
||||
|
||||
// A config written before `schedule` existed still works.
|
||||
// Empty or garbage must not stop the daemon.
|
||||
g.schedule = String::new();
|
||||
g.interval_mins = Some(45);
|
||||
assert_eq!(g.interval(), 45);
|
||||
|
||||
// Garbage must not stop the daemon.
|
||||
assert_eq!(g.interval(), 60);
|
||||
g.schedule = "whenever".into();
|
||||
assert_eq!(g.interval(), 45);
|
||||
g.interval_mins = None;
|
||||
assert_eq!(g.interval(), 60);
|
||||
}
|
||||
|
||||
@@ -467,7 +473,7 @@ mod tests {
|
||||
taken.insert("the-daily".to_string(), Feed {
|
||||
url: "u".into(), folder: None, group: None, media_types: None, schedule: None, keywords: vec![], allow_explicit: false,
|
||||
auto_download: true, max_new_per_check: None, username: None,
|
||||
password: None, password_env: None,
|
||||
password: None, password_env: None, category: None,
|
||||
});
|
||||
assert_eq!(unique_slug("The Daily", &taken), "the-daily-2");
|
||||
}
|
||||
@@ -487,6 +493,7 @@ mod tests {
|
||||
username: Some("ray".into()),
|
||||
password: Some("literal".into()),
|
||||
password_env: None,
|
||||
category: None,
|
||||
};
|
||||
assert_eq!(f.password().as_deref(), Some("literal"));
|
||||
|
||||
|
||||
@@ -223,7 +223,7 @@ enum Sniffed {
|
||||
/// 2008 and so always answered 'data'.
|
||||
async fn sniff(path: &Path) -> Result<Sniffed> {
|
||||
let head = read_head(path, 512).await?;
|
||||
if infer::is(&head, "torrent") || head.starts_with(b"d8:announce") || head.starts_with(b"d7:") {
|
||||
if head.starts_with(b"d8:announce") || head.starts_with(b"d7:") {
|
||||
return Ok(Sniffed::Torrent);
|
||||
}
|
||||
let text = String::from_utf8_lossy(&head);
|
||||
@@ -387,7 +387,7 @@ mod tests {
|
||||
let mut f = crate::config::Feed {
|
||||
url: "u".into(), folder: Some("Subscriptions/Some | Show".into()), group: None, media_types: None,
|
||||
schedule: None, keywords: vec![], allow_explicit: false, auto_download: true,
|
||||
max_new_per_check: None, username: None, password: None, password_env: None,
|
||||
max_new_per_check: None, username: None, password: None, password_env: None, category: None,
|
||||
};
|
||||
assert_eq!(folder_for(&cfg, "id", &f, None), "Subscriptions/Some - Show");
|
||||
|
||||
|
||||
538
src/feed.rs
538
src/feed.rs
@@ -11,6 +11,8 @@ pub struct ParsedFeed {
|
||||
pub title: Option<String>,
|
||||
pub ttl_mins: Option<u64>,
|
||||
pub image: Option<String>,
|
||||
/// The channel's first `<itunes:category>`, for the Directory's chips.
|
||||
pub category: Option<String>,
|
||||
pub entries: Vec<Entry>,
|
||||
}
|
||||
|
||||
@@ -87,6 +89,53 @@ pub async fn fetch(
|
||||
Ok(Fetched::Body { bytes, etag, last_modified })
|
||||
}
|
||||
|
||||
/// A stored `last_error`, translated into plain words for whoever subscribes: whose problem
|
||||
/// it is, and whether there is a new address to switch to.
|
||||
pub struct Failure {
|
||||
pub reason: &'static str,
|
||||
pub new_url: Option<String>,
|
||||
}
|
||||
|
||||
/// Reads a `last_error` the same way `set_feed_error` received it (`format!("{e:#}")` on the
|
||||
/// anyhow chain from `fetch` or `parse`) and says what it means, for the errors worth telling
|
||||
/// someone about. Everything else -- a timeout, a 5xx, a 429, a feed that is simply garbled --
|
||||
/// comes back `None`: transient by nature, or with nothing more useful to say than the raw
|
||||
/// text already shown once a feed is open.
|
||||
///
|
||||
/// ponytail: matches on the fixed strings this crate itself produces (`anyhow!("HTTP
|
||||
/// {status}")`, and `parse`'s "got a web page" and "the site sent") plus the substrings a DNS failure
|
||||
/// reliably contains. Fragile if reqwest's own wording changes; the fallback is just showing
|
||||
/// nothing extra, so a miss costs a clearer message, not a wrong one.
|
||||
pub fn explain_failure(msg: &str) -> Option<Failure> {
|
||||
if let Some(rest) = msg.strip_prefix("got a web page, not a feed") {
|
||||
let new_url = rest
|
||||
.strip_prefix("; it links ")
|
||||
.and_then(|r| r.strip_suffix(" as its feed"))
|
||||
.map(str::to_owned);
|
||||
return Some(Failure { reason: "The feed moved; this address now shows a web page.", new_url });
|
||||
}
|
||||
if msg.contains("the site sent ") {
|
||||
return Some(Failure { reason: "The site sent a message instead of the feed; the publisher has to fix it.", new_url: None });
|
||||
}
|
||||
let low = msg.to_ascii_lowercase();
|
||||
if low.contains("http 404") {
|
||||
return Some(Failure { reason: "The publisher took this feed down, or moved it.", new_url: None });
|
||||
}
|
||||
if low.contains("http 401") || low.contains("http 403") {
|
||||
return Some(Failure { reason: "The site refuses ipx's requests.", new_url: None });
|
||||
}
|
||||
if low.contains("http 402") {
|
||||
return Some(Failure { reason: "The feed now needs a paid plan.", new_url: None });
|
||||
}
|
||||
if low.contains("dns error")
|
||||
|| low.contains("failed to lookup address")
|
||||
|| low.contains("no address associated")
|
||||
{
|
||||
return Some(Failure { reason: "This address no longer resolves; the site is gone.", new_url: None });
|
||||
}
|
||||
None
|
||||
}
|
||||
|
||||
/// True when a body is an OPML document rather than a feed.
|
||||
///
|
||||
/// The original matched on the URL ending in ".opml" (iPXClass.py:34), which misses an
|
||||
@@ -117,17 +166,236 @@ pub fn opml_title(bytes: &[u8]) -> Option<String> {
|
||||
.filter(|t| !t.is_empty())
|
||||
}
|
||||
|
||||
/// The token and show of a Patreon feed link, or None for any other URL.
|
||||
///
|
||||
/// Patreon gives each patron one token per creator. With no show it stands for the creator,
|
||||
/// whose feed carries every show at once.
|
||||
fn patreon_parts(url: &str) -> Option<(String, Option<String>)> {
|
||||
let u = url::Url::parse(url).ok()?;
|
||||
if !matches!(u.host_str()?, "patreon.com" | "www.patreon.com") || !u.path().starts_with("/rss") {
|
||||
return None;
|
||||
}
|
||||
let param = |name: &str| u.query_pairs().find(|(k, _)| k == name).map(|(_, v)| v.into_owned());
|
||||
Some((param("auth")?, param("show")))
|
||||
}
|
||||
|
||||
/// A Patreon link naming a creator but no show.
|
||||
pub fn is_patreon_creator(url: &str) -> bool {
|
||||
matches!(patreon_parts(url), Some((_, None)))
|
||||
}
|
||||
|
||||
/// What was typed into Add feed, as a URL. A bare Patreon token is taken as its creator's
|
||||
/// feed, since the token alone says whose it is.
|
||||
pub fn expand_input(input: &str) -> String {
|
||||
let s = input.trim();
|
||||
let token = s.len() >= 20 && s.chars().all(|c| c.is_ascii_alphanumeric() || c == '-' || c == '_');
|
||||
if token { format!("https://www.patreon.com/rss?auth={s}") } else { s.to_owned() }
|
||||
}
|
||||
|
||||
/// Whether two URLs are the same feed. One Patreon show has several spellings -- by the
|
||||
/// creator's name, by number, or with no creator at all -- and the token and show are what
|
||||
/// identify it.
|
||||
pub fn same_feed(a: &str, b: &str) -> bool {
|
||||
a == b || patreon_parts(a).is_some_and(|p| Some(p) == patreon_parts(b))
|
||||
}
|
||||
|
||||
/// A Patreon creator's name and shows, each show as (title, feed URL).
|
||||
///
|
||||
/// ponytail: Patreon's own web API, undocumented, asked without signing in. If it changes,
|
||||
/// finding shows stops and the show feeds already found keep working. The documented API
|
||||
/// needs an OAuth client per install and does not list shows.
|
||||
pub async fn patreon_shows(
|
||||
client: &reqwest::Client,
|
||||
url: &str,
|
||||
) -> Result<(Option<String>, Vec<(String, String)>)> {
|
||||
// The creator feed names its campaign by number in its self link, a few hundred bytes in.
|
||||
// The whole feed runs to megabytes and Patreon ignores Range, so read until it turns up.
|
||||
let mut resp = client.get(url).send().await.context("connecting")?;
|
||||
if !resp.status().is_success() {
|
||||
return Err(anyhow!("Patreon refused the feed: HTTP {}", resp.status()));
|
||||
}
|
||||
let mut head = Vec::new();
|
||||
while patreon_campaign(&head).is_none() && head.len() < 64 * 1024 {
|
||||
let Some(chunk) = resp.chunk().await.context("reading the feed")? else { break };
|
||||
head.extend_from_slice(&chunk);
|
||||
}
|
||||
let campaign = patreon_campaign(&head)
|
||||
.ok_or_else(|| anyhow!("the Patreon feed does not say whose it is"))?;
|
||||
|
||||
let api = format!(
|
||||
"https://www.patreon.com/api/campaigns/{campaign}\
|
||||
?include=shows&fields%5Bcampaign%5D=name&fields%5Bcollection%5D=title"
|
||||
);
|
||||
let resp = client.get(api).send().await.context("asking Patreon for the shows")?;
|
||||
if !resp.status().is_success() {
|
||||
return Err(anyhow!("Patreon would not list the shows: HTTP {}", resp.status()));
|
||||
}
|
||||
let (name, shows) = parse_patreon_shows(&resp.bytes().await.context("reading the shows")?)?;
|
||||
Ok((name, shows.into_iter().map(|(id, title)| (title, format!("{url}&show={id}"))).collect()))
|
||||
}
|
||||
|
||||
/// The campaign number in the start of a Patreon feed.
|
||||
fn patreon_campaign(head: &[u8]) -> Option<String> {
|
||||
let text = String::from_utf8_lossy(head);
|
||||
text.match_indices("patreon.com/rss/").find_map(|(i, m)| {
|
||||
let id: String = text[i + m.len()..].chars().take_while(char::is_ascii_digit).collect();
|
||||
(!id.is_empty()).then_some(id)
|
||||
})
|
||||
}
|
||||
|
||||
/// A campaign's name and its shows as (id, title), from Patreon's JSON:API answer.
|
||||
fn parse_patreon_shows(json: &[u8]) -> Result<(Option<String>, Vec<(String, String)>)> {
|
||||
let v: serde_json::Value = serde_json::from_slice(json).context("Patreon's answer is not JSON")?;
|
||||
// Missing is not the same as none. Read as no shows, the creator feed would be scanned as
|
||||
// a plain feed, claim every show's files, and leave the shows empty once the list returned.
|
||||
let ids = v["data"]["relationships"]["shows"]["data"]
|
||||
.as_array()
|
||||
.ok_or_else(|| anyhow!("Patreon's answer does not list the shows"))?;
|
||||
let title = |id: &str| -> Option<String> {
|
||||
let show = v["included"].as_array()?.iter().find(|x| x["type"] == "collection" && x["id"] == id)?;
|
||||
show["attributes"]["title"].as_str().map(|t| t.trim().to_owned())
|
||||
};
|
||||
let shows = ids
|
||||
.iter()
|
||||
.filter_map(|s| s["id"].as_str())
|
||||
.map(|id| (id.to_owned(), title(id).unwrap_or_else(|| format!("Show {id}"))))
|
||||
.collect();
|
||||
Ok((v["data"]["attributes"]["name"].as_str().map(str::to_owned), shows))
|
||||
}
|
||||
|
||||
/// RSS first, then Atom -- the same split the original made on `parsedFeed.version`.
|
||||
pub fn parse(bytes: &[u8]) -> Result<ParsedFeed> {
|
||||
match rss::Channel::read_from(bytes) {
|
||||
Ok(ch) => Ok(from_rss(ch, bytes)),
|
||||
Err(rss_err) => match atom_syndication::Feed::read_from(bytes) {
|
||||
Ok(feed) => Ok(from_atom(feed)),
|
||||
Err(atom_err) => Err(anyhow!("not RSS ({rss_err}) and not Atom ({atom_err})")),
|
||||
Err(atom_err) => {
|
||||
if let Some(said) = plain_text(bytes) {
|
||||
return Err(anyhow!("the site sent {said} instead of a feed"));
|
||||
}
|
||||
// Some publishers (kcpw, feedland) write a bare "&" in a URL instead of
|
||||
// "&". Strict XML parsers refuse it; browsers don't. Retry once with
|
||||
// every offending "&" escaped rather than fail outright.
|
||||
let escaped = escape_bare_ampersands(bytes);
|
||||
if escaped != bytes {
|
||||
if let Ok(ch) = rss::Channel::read_from(escaped.as_slice()) {
|
||||
return Ok(from_rss(ch, &escaped));
|
||||
}
|
||||
if let Ok(feed) = atom_syndication::Feed::read_from(escaped.as_slice()) {
|
||||
return Ok(from_atom(feed));
|
||||
}
|
||||
}
|
||||
Err(match alternate_feed_link(bytes) {
|
||||
Some(href) if looks_like_html(bytes) => {
|
||||
anyhow!("got a web page, not a feed; it links {href} as its feed")
|
||||
}
|
||||
None if looks_like_html(bytes) => anyhow!("got a web page, not a feed"),
|
||||
_ => anyhow!("not RSS ({rss_err}) and not Atom ({atom_err})"),
|
||||
})
|
||||
}
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether a body is a web page rather than a feed: most of the errors traced back to a feed
|
||||
/// that moved or a domain that lapsed, with the old URL now serving the site instead (or a
|
||||
/// redirect to it). `is_opml` already sniffs the other "not actually a feed" case.
|
||||
fn looks_like_html(bytes: &[u8]) -> bool {
|
||||
let head = String::from_utf8_lossy(&bytes[..bytes.len().min(2048)]).to_lowercase();
|
||||
head.contains("<!doctype html") || head.contains("<html")
|
||||
}
|
||||
|
||||
/// What a site sent when it sent a sentence instead of markup. doghouse's feed answered 200 with
|
||||
/// "Unable to establish a DB connection", and the two parsers' errors about end of input buried
|
||||
/// it. Anything starting with `<` is markup, however broken, and keeps the parsers' errors.
|
||||
fn plain_text(bytes: &[u8]) -> Option<String> {
|
||||
let head = String::from_utf8_lossy(&bytes[..bytes.len().min(512)]);
|
||||
let text = head.trim_start_matches(|c: char| c.is_whitespace() || c == '\u{feff}');
|
||||
if text.starts_with('<') {
|
||||
return None;
|
||||
}
|
||||
let line = text.lines().next().unwrap_or("").trim_end();
|
||||
if line.is_empty() {
|
||||
return Some("an empty reply".into());
|
||||
}
|
||||
let mut said: String = line.chars().take(80).collect();
|
||||
if said.len() < line.len() {
|
||||
said.push('…');
|
||||
}
|
||||
Some(format!("\"{said}\""))
|
||||
}
|
||||
|
||||
/// The feed a web page names as its own via `<link rel="alternate" type="application/rss+xml"
|
||||
/// href="...">` (or the Atom equivalent) -- how the new address was found for om.co, ms.now,
|
||||
/// Letters of Note, the Daily Dot, Hell Gate, The Frame Lab and Daily Kos.
|
||||
fn alternate_feed_link(bytes: &[u8]) -> Option<String> {
|
||||
let text = String::from_utf8_lossy(bytes);
|
||||
let lower = text.to_lowercase();
|
||||
let mut pos = 0;
|
||||
while let Some(rel) = lower[pos..].find("<link") {
|
||||
let start = pos + rel;
|
||||
let Some(end) = lower[start..].find('>').map(|e| start + e) else { break };
|
||||
pos = end + 1;
|
||||
let tag = &text[start..end];
|
||||
let tag_lower = &lower[start..end];
|
||||
let is_alternate = tag_lower.contains("rel=\"alternate\"") || tag_lower.contains("rel='alternate'");
|
||||
let is_feed_type = tag_lower.contains("rss+xml") || tag_lower.contains("atom+xml");
|
||||
if is_alternate && is_feed_type
|
||||
&& let Some(href) = tag_attr(tag, "href")
|
||||
{
|
||||
return Some(href);
|
||||
}
|
||||
}
|
||||
None
|
||||
}
|
||||
|
||||
/// The value of one attribute in an HTML/XML start tag, however it is quoted.
|
||||
fn tag_attr(tag: &str, name: &str) -> Option<String> {
|
||||
let key = format!("{name}=");
|
||||
let idx = tag.to_lowercase().find(&key)?;
|
||||
let after = &tag[idx + key.len()..];
|
||||
let quote = after.chars().next()?;
|
||||
if quote != '"' && quote != '\'' {
|
||||
return None;
|
||||
}
|
||||
let rest = &after[1..];
|
||||
let close = rest.find(quote)?;
|
||||
Some(rest[..close].trim().to_owned())
|
||||
}
|
||||
|
||||
/// Escapes every `&` that does not already start a recognized XML entity
|
||||
/// (`&`, `<`, `>`, `"`, `'`, or a numeric reference like `'`).
|
||||
fn escape_bare_ampersands(bytes: &[u8]) -> Vec<u8> {
|
||||
fn is_entity_start(rest: &[u8]) -> bool {
|
||||
for named in [&b"amp;"[..], b"lt;", b"gt;", b"quot;", b"apos;"] {
|
||||
if rest.starts_with(named) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
let digits = if rest.starts_with(b"#x") || rest.starts_with(b"#X") {
|
||||
&rest[2..]
|
||||
} else if rest.starts_with(b"#") {
|
||||
&rest[1..]
|
||||
} else {
|
||||
return false;
|
||||
};
|
||||
let len = digits.iter().take_while(|b| b.is_ascii_alphanumeric()).count();
|
||||
len > 0 && digits.get(len) == Some(&b';')
|
||||
}
|
||||
|
||||
let mut out = Vec::with_capacity(bytes.len());
|
||||
let mut i = 0;
|
||||
while i < bytes.len() {
|
||||
if bytes[i] == b'&' && !is_entity_start(&bytes[i + 1..]) {
|
||||
out.extend_from_slice(b"&");
|
||||
} else {
|
||||
out.push(bytes[i]);
|
||||
}
|
||||
i += 1;
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Every `<enclosure>` of every `<item>`, in document order.
|
||||
///
|
||||
/// The `rss` crate models an item as having at most one enclosure -- which is what RSS 2.0
|
||||
@@ -212,7 +480,7 @@ fn from_rss(ch: rss::Channel, bytes: &[u8]) -> ParsedFeed {
|
||||
.filter_map(|(idx, item)| {
|
||||
// Straight from the XML, so an item with several keeps all of them. Falls
|
||||
// back to the parsed one if the scan and the parser disagree on item count.
|
||||
let enclosures: Vec<Enclosure> = per_item.get(idx).cloned().unwrap_or_else(|| {
|
||||
let mut enclosures: Vec<Enclosure> = per_item.get(idx).cloned().unwrap_or_else(|| {
|
||||
item.enclosure()
|
||||
.into_iter()
|
||||
.map(|e| Enclosure {
|
||||
@@ -223,6 +491,7 @@ fn from_rss(ch: rss::Channel, bytes: &[u8]) -> ParsedFeed {
|
||||
.filter(|e| !e.url.is_empty())
|
||||
.collect()
|
||||
});
|
||||
drop_player_repeats(&mut enclosures);
|
||||
|
||||
let guid = pick_guid(
|
||||
item.guid().map(|g| g.value()),
|
||||
@@ -239,11 +508,11 @@ fn from_rss(ch: rss::Channel, bytes: &[u8]) -> ParsedFeed {
|
||||
|
||||
Some(Entry {
|
||||
guid,
|
||||
title: non_empty(item.title()),
|
||||
title: title_text(item.title()),
|
||||
link: non_empty(item.link()),
|
||||
published: item.pub_date().and_then(parse_date),
|
||||
// Content wins over description, as __getEntries preferred entry.content.
|
||||
description: non_empty(item.content()).or_else(|| non_empty(item.description())),
|
||||
description: body(item.content(), item.description()),
|
||||
categories: item
|
||||
.categories()
|
||||
.iter()
|
||||
@@ -261,7 +530,7 @@ fn from_rss(ch: rss::Channel, bytes: &[u8]) -> ParsedFeed {
|
||||
.collect();
|
||||
|
||||
ParsedFeed {
|
||||
title: non_empty(Some(ch.title())),
|
||||
title: title_text(Some(ch.title())),
|
||||
ttl_mins: ch.ttl().and_then(|t| t.trim().parse().ok()),
|
||||
// itunes:image is the square artwork; <image><url> is the older, often smaller one.
|
||||
image: ch
|
||||
@@ -269,6 +538,14 @@ fn from_rss(ch: rss::Channel, bytes: &[u8]) -> ParsedFeed {
|
||||
.and_then(|i| i.image())
|
||||
.map(str::to_owned)
|
||||
.or_else(|| ch.image().map(|i| i.url().to_owned())),
|
||||
// Only the iTunes one: Apple's list is fixed, while a plain <category> is freeform and
|
||||
// would fill the Directory with one-off tags. The subcategory where there is one: Apple
|
||||
// files every tabletop and gaming show under Leisure, which says little; Games says it.
|
||||
category: ch
|
||||
.itunes_ext()
|
||||
.and_then(|i| i.categories().first())
|
||||
.map(|c| c.subcategory().filter(|s| !s.text().trim().is_empty()).unwrap_or(c))
|
||||
.and_then(|c| non_empty(Some(c.text().trim()))),
|
||||
entries,
|
||||
}
|
||||
}
|
||||
@@ -279,7 +556,7 @@ fn from_atom(feed: atom_syndication::Feed) -> ParsedFeed {
|
||||
.iter()
|
||||
.filter_map(|e| {
|
||||
// Atom carries enclosures as <link rel="enclosure">.
|
||||
let enclosures: Vec<Enclosure> = e
|
||||
let mut enclosures: Vec<Enclosure> = e
|
||||
.links()
|
||||
.iter()
|
||||
.filter(|l| l.rel() == "enclosure")
|
||||
@@ -290,6 +567,7 @@ fn from_atom(feed: atom_syndication::Feed) -> ParsedFeed {
|
||||
})
|
||||
.filter(|e| !e.url.is_empty())
|
||||
.collect();
|
||||
drop_player_repeats(&mut enclosures);
|
||||
|
||||
let alt = e
|
||||
.links()
|
||||
@@ -306,14 +584,10 @@ fn from_atom(feed: atom_syndication::Feed) -> ParsedFeed {
|
||||
|
||||
Some(Entry {
|
||||
guid,
|
||||
title: non_empty(Some(e.title().as_str())),
|
||||
title: title_text(Some(e.title().as_str())),
|
||||
link: alt.map(str::to_owned),
|
||||
published: e.published().or(Some(e.updated())).map(|d| d.timestamp()),
|
||||
description: e
|
||||
.content()
|
||||
.and_then(|c| c.value())
|
||||
.or_else(|| e.summary().map(|s| s.as_str()))
|
||||
.map(str::to_owned),
|
||||
description: body(e.content().and_then(|c| c.value()), e.summary().map(|s| s.as_str())),
|
||||
categories: e.categories().iter().map(|c| c.term().to_owned()).collect(),
|
||||
explicit: false,
|
||||
image: None,
|
||||
@@ -326,9 +600,10 @@ fn from_atom(feed: atom_syndication::Feed) -> ParsedFeed {
|
||||
.collect();
|
||||
|
||||
ParsedFeed {
|
||||
title: non_empty(Some(feed.title().as_str())),
|
||||
title: title_text(Some(feed.title().as_str())),
|
||||
ttl_mins: None,
|
||||
image: feed.logo().or_else(|| feed.icon()).map(str::to_owned),
|
||||
category: None,
|
||||
entries,
|
||||
}
|
||||
}
|
||||
@@ -358,6 +633,83 @@ fn non_empty(s: Option<&str>) -> Option<String> {
|
||||
s.map(str::trim).filter(|s| !s.is_empty()).map(str::to_owned)
|
||||
}
|
||||
|
||||
/// WordPress numbers each audio player on a page by adding `?_=N` to its file's URL, so a post
|
||||
/// that embeds the file it encloses lists the same file twice: Rands in Repose's "The Promotion
|
||||
/// Paradox" was downloaded twice and offered two play buttons for one mp3. The first stays.
|
||||
fn drop_player_repeats(encs: &mut Vec<Enclosure>) {
|
||||
let mut seen = std::collections::HashSet::new();
|
||||
encs.retain(|e| seen.insert(same_file_key(&e.url)));
|
||||
}
|
||||
|
||||
/// An enclosure URL without WordPress's player number, for telling repeats of one file apart
|
||||
/// from different files.
|
||||
pub fn same_file_key(url: &str) -> String {
|
||||
let Ok(mut u) = url::Url::parse(url) else { return url.to_owned() };
|
||||
let kept: Vec<(String, String)> = u
|
||||
.query_pairs()
|
||||
.filter(|(k, v)| !(k == "_" && !v.is_empty() && v.bytes().all(|b| b.is_ascii_digit())))
|
||||
.map(|(k, v)| (k.into_owned(), v.into_owned()))
|
||||
.collect();
|
||||
if kept.is_empty() {
|
||||
u.set_query(None);
|
||||
} else {
|
||||
u.query_pairs_mut().clear().extend_pairs(kept);
|
||||
}
|
||||
u.to_string()
|
||||
}
|
||||
|
||||
/// A title as plain text. An Atom title of `type="html"`, or an RSS one in CDATA, comes through
|
||||
/// the XML parser with its HTML entities intact: The Verge's "Meta’s" reached the page as
|
||||
/// typed. Decoded one entity at a time, so an `&` that starts none, as in "Q&A", stays as it is
|
||||
/// instead of failing the whole title.
|
||||
fn title_text(s: Option<&str>) -> Option<String> {
|
||||
let s = non_empty(s)?;
|
||||
let mut out = String::with_capacity(s.len());
|
||||
let mut rest = s.as_str();
|
||||
while let Some(at) = rest.find('&') {
|
||||
out.push_str(&rest[..at]);
|
||||
rest = &rest[at..];
|
||||
let len = 1 + rest[1..]
|
||||
.find(|c: char| !(c.is_ascii_alphanumeric() || c == '#'))
|
||||
.unwrap_or(rest.len() - 1);
|
||||
let decoded = rest[len..]
|
||||
.starts_with(';')
|
||||
.then(|| quick_xml::escape::unescape_with(&rest[..=len], quick_xml::escape::resolve_html5_entity).ok())
|
||||
.flatten();
|
||||
match decoded {
|
||||
Some(v) => {
|
||||
out.push_str(&v);
|
||||
rest = &rest[len + 1..];
|
||||
}
|
||||
None => {
|
||||
out.push('&');
|
||||
rest = &rest[1..];
|
||||
}
|
||||
}
|
||||
}
|
||||
out.push_str(rest);
|
||||
non_empty(Some(&out))
|
||||
}
|
||||
|
||||
/// An item's show notes: its full body when that is whole, else its description.
|
||||
///
|
||||
/// libsyn served Daily Meditation Podcast's `content:encoded` cut at the `>` inside a class name
|
||||
/// pasted from a web app (`[&:has([data-writing-block])>*]:pointer-events-auto`), so the body
|
||||
/// began halfway through a tag and the page showed the rest of the tag as text. The same item's
|
||||
/// `description` was whole. With no description to fall back on, a damaged body beats none.
|
||||
fn body(content: Option<&str>, description: Option<&str>) -> Option<String> {
|
||||
non_empty(content)
|
||||
.filter(|c| !starts_mid_tag(c))
|
||||
.or_else(|| non_empty(description))
|
||||
.or_else(|| non_empty(content))
|
||||
}
|
||||
|
||||
/// Text that closes an attribute list (`">`) before any tag has opened is the tail of a tag whose
|
||||
/// start was cut off.
|
||||
fn starts_mid_tag(html: &str) -> bool {
|
||||
html[..html.find('<').unwrap_or(html.len())].contains("\">")
|
||||
}
|
||||
|
||||
/// The picture to show beside an item, in order of how deliberate it is:
|
||||
/// `itunes:image`, then Media RSS `media:thumbnail`, then a `media:content` that is an
|
||||
/// image, and finally an image enclosure -- which is how a blog's article picture arrives
|
||||
@@ -427,6 +779,7 @@ mod tests {
|
||||
|
||||
assert_eq!(feed.title.as_deref(), Some("Test Cast"));
|
||||
assert_eq!(feed.ttl_mins, Some(45));
|
||||
assert_eq!(feed.category.as_deref(), Some("Podcasting"), "the first, by its subcategory");
|
||||
assert_eq!(feed.entries.len(), 3);
|
||||
|
||||
let ep = &feed.entries[0];
|
||||
@@ -451,6 +804,49 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_file_wordpress_lists_twice_is_one_enclosure() {
|
||||
let xml = br#"<?xml version="1.0"?><rss version="2.0"><channel><title>R</title>
|
||||
<item><title>The Promotion Paradox</title><guid>p</guid>
|
||||
<enclosure url="https://x/ep.mp3" length="1" type="audio/mpeg"/>
|
||||
<enclosure url="https://x/ep.mp3?_=2" length="1" type="audio/mpeg"/>
|
||||
<enclosure url="https://x/other.mp3?_=3&key=k" length="1" type="audio/mpeg"/>
|
||||
</item></channel></rss>"#;
|
||||
let urls: Vec<String> =
|
||||
parse(xml).unwrap().entries[0].enclosures.iter().map(|e| e.url.clone()).collect();
|
||||
assert_eq!(urls, ["https://x/ep.mp3", "https://x/other.mp3?_=3&key=k"], "the repeat goes, a different file stays");
|
||||
assert_eq!(same_file_key("https://x/a.mp3?key=k&_=2"), same_file_key("https://x/a.mp3?key=k"));
|
||||
assert_ne!(same_file_key("https://x/a.mp3?_=x"), same_file_key("https://x/a.mp3"), "only a number");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn titles_are_read_as_text_not_html() {
|
||||
// The Verge: an Atom title of type="html", its entity inside CDATA.
|
||||
let xml = br#"<?xml version="1.0"?>
|
||||
<feed xmlns="http://www.w3.org/2005/Atom"><title type="text">V</title><id>v</id>
|
||||
<updated>2026-09-15T00:00:00Z</updated>
|
||||
<entry><title type="html"><![CDATA[Meta’s new One]]></title><id>e1</id>
|
||||
<updated>2026-09-15T00:00:00Z</updated></entry></feed>"#;
|
||||
assert_eq!(parse(xml).unwrap().entries[0].title.as_deref(), Some("Meta\u{2019}s new One"));
|
||||
// HTML names as well as numbers; a bare `&` and an unknown name are left as they are.
|
||||
assert_eq!(
|
||||
title_text(Some("Peña & “Q&A” &bogus; AT&T;")).as_deref(),
|
||||
Some("Pe\u{f1}a & \u{201c}Q&A\u{201d} &bogus; AT&T;")
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_body_cut_off_mid_tag_gives_way_to_the_description() {
|
||||
// How libsyn served Daily Meditation Podcast #3477: content:encoded began inside a tag.
|
||||
let cut = r#"*]:pointer-events-auto R6Vx5W_threadScrollVars" dir="auto" data-turn="assistant"> <p>What if</p>"#;
|
||||
let whole = r#"<div class="[&:has([data-writing-block])>*]:pointer-events-auto"><p>What if</p></div>"#;
|
||||
assert_eq!(body(Some(cut), Some(whole)).as_deref(), Some(whole));
|
||||
assert_eq!(body(Some("<p>Notes</p>"), Some("Summary")).as_deref(), Some("<p>Notes</p>"), "a whole body wins");
|
||||
assert_eq!(body(Some("Plain notes, no tags."), Some("Summary")).as_deref(), Some("Plain notes, no tags."));
|
||||
assert_eq!(body(Some(cut), None).as_deref(), Some(cut), "a damaged body beats none");
|
||||
assert_eq!(body(None, Some("Summary")).as_deref(), Some("Summary"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn feed_level_explicit_overrides_entries() {
|
||||
let xml = br#"<?xml version="1.0"?>
|
||||
@@ -467,6 +863,62 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn explain_failure_translates_the_errors_the_ui_should_flag() {
|
||||
assert_eq!(
|
||||
explain_failure("HTTP 404 Not Found").unwrap().reason,
|
||||
"The publisher took this feed down, or moved it."
|
||||
);
|
||||
assert_eq!(explain_failure("HTTP 401 Unauthorized").unwrap().reason, "The site refuses ipx's requests.");
|
||||
assert_eq!(explain_failure("HTTP 403 Forbidden").unwrap().reason, "The site refuses ipx's requests.");
|
||||
assert_eq!(explain_failure("HTTP 402 Payment Required").unwrap().reason, "The feed now needs a paid plan.");
|
||||
let dns = explain_failure("connecting: dns error: failed to lookup address information").unwrap();
|
||||
assert_eq!(dns.reason, "This address no longer resolves; the site is gone.");
|
||||
let moved = explain_failure("got a web page, not a feed; it links https://x/feed as its feed").unwrap();
|
||||
assert_eq!(moved.new_url.as_deref(), Some("https://x/feed"));
|
||||
assert!(explain_failure("got a web page, not a feed").unwrap().new_url.is_none());
|
||||
let down = explain_failure("the site sent \"Unable to establish a DB connection\" instead of a feed").unwrap();
|
||||
assert_eq!(down.reason, "The site sent a message instead of the feed; the publisher has to fix it.");
|
||||
for transient in [
|
||||
"HTTP 500 Internal Server Error",
|
||||
"HTTP 429 Too Many Requests",
|
||||
"operation timed out",
|
||||
"not RSS (reached end of input without finding a complete channel) and not Atom (unexpected end of input)",
|
||||
] {
|
||||
assert!(explain_failure(transient).is_none(), "{transient} must not be flagged");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_web_page_says_so_and_names_the_feed_it_links() {
|
||||
let html = br#"<!doctype html><html><head>
|
||||
<link rel="alternate" type="application/rss+xml" href="https://x.example/feed">
|
||||
</head><body>not a feed</body></html>"#;
|
||||
let err = parse(html).unwrap_err().to_string();
|
||||
assert_eq!(err, "got a web page, not a feed; it links https://x.example/feed as its feed");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_web_page_with_no_feed_link_still_says_so() {
|
||||
let html = b"<!doctype html><html><body>moved</body></html>";
|
||||
assert_eq!(parse(html).unwrap_err().to_string(), "got a web page, not a feed");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn malformed_xml_gets_the_original_parser_errors() {
|
||||
let err = parse(b"<rss><channel><title>cut off").unwrap_err().to_string();
|
||||
assert!(err.starts_with("not RSS ("), "{err}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_body_with_no_markup_says_what_the_site_sent() {
|
||||
let err = parse(b"\xef\xbb\xbf\r\n Unable to establish a DB connection\nmore").unwrap_err().to_string();
|
||||
assert_eq!(err, "the site sent \"Unable to establish a DB connection\" instead of a feed");
|
||||
let long = parse("x".repeat(200).as_bytes()).unwrap_err().to_string();
|
||||
assert_eq!(long, format!("the site sent \"{}…\" instead of a feed", "x".repeat(80)));
|
||||
assert_eq!(parse(b" \n").unwrap_err().to_string(), "the site sent an empty reply instead of a feed");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parses_atom_enclosure_links() {
|
||||
let bytes = include_bytes!("../tests/data/atom.xml");
|
||||
@@ -489,6 +941,27 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_bare_ampersand_in_a_link_is_repaired_and_parsed() {
|
||||
// kcpw.org: <link>https://kcpw.org/?post_type=post&p=125715</link> -- a bare "&"
|
||||
// that strict XML rejects but browsers accept.
|
||||
let xml = br#"<?xml version="1.0"?>
|
||||
<rss version="2.0"><channel><title>X</title><link>https://x</link><description>d</description>
|
||||
<item><title>a</title><guid>g1</guid>
|
||||
<link>https://kcpw.org/?post_type=post&p=125715</link>
|
||||
<enclosure url="https://x/a.mp3?a=1&b=2" length="1" type="audio/mpeg"/></item>
|
||||
</channel></rss>"#;
|
||||
let feed = parse(xml).unwrap();
|
||||
assert_eq!(feed.entries[0].link.as_deref(), Some("https://kcpw.org/?post_type=post&p=125715"));
|
||||
assert_eq!(feed.entries[0].enclosures[0].url, "https://x/a.mp3?a=1&b=2");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn escape_bare_ampersands_leaves_real_entities_alone() {
|
||||
let out = escape_bare_ampersands(b"a&b <x> ' / c&d");
|
||||
assert_eq!(out, b"a&b <x> ' / c&d");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_rss_title_always_wins_and_episode_numbers_stay_metadata() {
|
||||
// Some feeds set a different itunes:title. The displayed title is always the RSS
|
||||
@@ -550,6 +1023,45 @@ mod tests {
|
||||
assert!(!is_opml(include_bytes!("../tests/data/atom.xml")));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_patreon_creator_is_a_list_of_its_shows() {
|
||||
let tok = "AbCdEfGhIjKlMnOpQrStUvWxYz012_-9";
|
||||
assert_eq!(expand_input(&format!(" {tok} ")), format!("https://www.patreon.com/rss?auth={tok}"));
|
||||
assert_eq!(expand_input("https://example.com/rss"), "https://example.com/rss");
|
||||
|
||||
assert!(is_patreon_creator(&format!("https://www.patreon.com/rss/glasscannon?auth={tok}")));
|
||||
assert!(is_patreon_creator(&format!("https://www.patreon.com/rss?auth={tok}")));
|
||||
assert!(!is_patreon_creator(&format!("https://www.patreon.com/rss/x?auth={tok}&show=1")), "one show is a feed");
|
||||
assert!(!is_patreon_creator(&format!("https://example.com/rss?auth={tok}")));
|
||||
|
||||
// The show you already have by name is the one a bare token would add by number.
|
||||
assert!(same_feed(
|
||||
&format!("https://www.patreon.com/rss/glasscannon?auth={tok}&show=2073588"),
|
||||
&format!("https://www.patreon.com/rss?auth={tok}&show=2073588"),
|
||||
));
|
||||
assert!(!same_feed(
|
||||
&format!("https://www.patreon.com/rss?auth={tok}&show=1"),
|
||||
&format!("https://www.patreon.com/rss?auth={tok}&show=2"),
|
||||
));
|
||||
|
||||
// The self link carries the campaign by number, whichever spelling was asked for.
|
||||
let head = br#"<rss><channel><link>https://www.patreon.com/glasscannon</link>
|
||||
<atom:link href="https://www.patreon.com/rss/369921?auth=t" rel="self"/>"#;
|
||||
assert_eq!(patreon_campaign(head).as_deref(), Some("369921"));
|
||||
assert_eq!(patreon_campaign(b"<rss><channel><title>T"), None);
|
||||
|
||||
let json = br#"{"data":{"id":"369921","type":"campaign","attributes":{"name":"The Glass Cannon Network"},
|
||||
"relationships":{"shows":{"data":[{"id":"2073588","type":"collection"},{"id":"2073636","type":"collection"}]}}},
|
||||
"included":[{"id":"2073588","type":"collection","attributes":{"title":"Get in the Trunk "}},
|
||||
{"id":"2073636","type":"collection","attributes":{"title":"Shadowdark"}}]}"#;
|
||||
let (name, shows) = parse_patreon_shows(json).unwrap();
|
||||
assert_eq!(name.as_deref(), Some("The Glass Cannon Network"));
|
||||
assert_eq!(shows, [("2073588".into(), "Get in the Trunk".into()), ("2073636".into(), "Shadowdark".into())]);
|
||||
|
||||
// An answer that stops naming the shows is an error, never "this creator has none".
|
||||
assert!(parse_patreon_shows(br#"{"data":{"attributes":{"name":"X"}}}"#).is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_item_may_carry_several_enclosures() {
|
||||
// The rss crate keeps only one per item -- the last -- so these come from the XML.
|
||||
|
||||
57
src/ipc.rs
57
src/ipc.rs
@@ -172,11 +172,18 @@ pub async fn daemon_is_live(path: &Path) -> bool {
|
||||
UnixStream::connect(path).await.is_ok()
|
||||
}
|
||||
|
||||
/// Answers `status` for the socket, without the worker. The worker runs one job at a time, and a
|
||||
/// healthcheck left waiting behind a scan or a long download timed out and called a busy daemon
|
||||
/// dead. The answer goes to the client that asked and no one else: broadcast, it ended any
|
||||
/// `ipx fetch` that was watching a scan, since `status` is a terminal event.
|
||||
pub type StatusFn = std::sync::Arc<dyn Fn() -> Event + Send + Sync>;
|
||||
|
||||
/// Accepts connections, feeding commands to `cmds` and events from `events` back out.
|
||||
pub async fn serve(
|
||||
path: PathBuf,
|
||||
events: broadcast::Sender<Event>,
|
||||
cmds: mpsc::Sender<Command>,
|
||||
status: StatusFn,
|
||||
) -> Result<()> {
|
||||
// A socket file left by a crashed daemon would block the bind; a live one was already
|
||||
// rejected by the caller's daemon_is_live() check.
|
||||
@@ -195,8 +202,9 @@ pub async fn serve(
|
||||
let (stream, _) = listener.accept().await?;
|
||||
let rx = events.subscribe();
|
||||
let cmds = cmds.clone();
|
||||
let status = status.clone();
|
||||
tokio::spawn(async move {
|
||||
if let Err(e) = handle(stream, rx, cmds).await {
|
||||
if let Err(e) = handle(stream, rx, cmds, status).await {
|
||||
tracing::debug!(error = %e, "client gone");
|
||||
}
|
||||
});
|
||||
@@ -207,12 +215,21 @@ async fn handle(
|
||||
stream: UnixStream,
|
||||
mut rx: broadcast::Receiver<Event>,
|
||||
cmds: mpsc::Sender<Command>,
|
||||
status: StatusFn,
|
||||
) -> Result<()> {
|
||||
let (read, mut write) = stream.into_split();
|
||||
|
||||
// Events out.
|
||||
// Events out: everything broadcast, and the answers meant for this client alone.
|
||||
let (reply, mut replies) = mpsc::channel::<Event>(4);
|
||||
let writer = tokio::spawn(async move {
|
||||
while let Ok(ev) = rx.recv().await {
|
||||
loop {
|
||||
let ev = tokio::select! {
|
||||
Some(ev) = replies.recv() => ev,
|
||||
got = rx.recv() => match got {
|
||||
Ok(ev) => ev,
|
||||
Err(_) => break,
|
||||
},
|
||||
};
|
||||
let mut line = serde_json::to_string(&ev).unwrap_or_default();
|
||||
line.push('\n');
|
||||
if write.write_all(line.as_bytes()).await.is_err() {
|
||||
@@ -229,6 +246,15 @@ async fn handle(
|
||||
continue;
|
||||
}
|
||||
match serde_json::from_str::<Command>(line) {
|
||||
// Answered here, not queued behind whatever the worker is on: see StatusFn.
|
||||
Ok(Command::Status) => {
|
||||
tracing::info!(target: "ipx::io", "-> {line}");
|
||||
let ev = status();
|
||||
if let Ok(json) = serde_json::to_string(&ev) {
|
||||
tracing::info!(target: "ipx::io", "<- {json}");
|
||||
}
|
||||
let _ = reply.send(ev).await;
|
||||
}
|
||||
Ok(cmd) => {
|
||||
if cmds.send(cmd).await.is_err() {
|
||||
break; // Worker is gone; so are we.
|
||||
@@ -329,4 +355,29 @@ mod tests {
|
||||
}
|
||||
.is_terminal());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn status_is_answered_while_the_worker_is_busy() {
|
||||
// The queue is full and nobody drains it, as when the worker is deep in a long download:
|
||||
// anything sent to it would wait for ever.
|
||||
let (cmds, _worker) = mpsc::channel::<Command>(1);
|
||||
cmds.send(Command::Reap { dry_run: true }).await.unwrap();
|
||||
let (events, _) = broadcast::channel::<Event>(8);
|
||||
// Another client, watching a scan: it must not be handed someone else's answer, which
|
||||
// would end its session.
|
||||
let mut watcher = events.subscribe();
|
||||
let status: StatusFn = std::sync::Arc::new(|| Event::Status { feeds: 1, pending: 2, downloaded: 3 });
|
||||
let (client, server) = UnixStream::pair().unwrap();
|
||||
tokio::spawn(handle(server, events.subscribe(), cmds, status));
|
||||
|
||||
let (read, mut write) = client.into_split();
|
||||
write.write_all(b"{\"cmd\":\"status\"}\n").await.unwrap();
|
||||
let line = tokio::time::timeout(std::time::Duration::from_secs(2), BufReader::new(read).lines().next_line())
|
||||
.await
|
||||
.expect("status waited behind the worker")
|
||||
.unwrap()
|
||||
.unwrap();
|
||||
assert!(line.contains(r#""ev":"status""#), "{line}");
|
||||
assert!(watcher.try_recv().is_err(), "the answer went to every client, not just the one asking");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -111,18 +111,11 @@ impl Visit for Collect {
|
||||
fn record_debug(&mut self, field: &Field, value: &dyn std::fmt::Debug) {
|
||||
self.add(field, format!("{value:?}"));
|
||||
}
|
||||
// Numbers and bools reach record_debug through the trait's defaults, which prints them the
|
||||
// same way. A string would print quoted there, hence its own method.
|
||||
fn record_str(&mut self, field: &Field, value: &str) {
|
||||
self.add(field, value.to_owned());
|
||||
}
|
||||
fn record_i64(&mut self, field: &Field, value: i64) {
|
||||
self.add(field, value.to_string());
|
||||
}
|
||||
fn record_u64(&mut self, field: &Field, value: u64) {
|
||||
self.add(field, value.to_string());
|
||||
}
|
||||
fn record_bool(&mut self, field: &Field, value: bool) {
|
||||
self.add(field, value.to_string());
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
|
||||
341
src/main.rs
341
src/main.rs
@@ -99,6 +99,9 @@ enum UserCmd {
|
||||
Passwd { name: String },
|
||||
/// Delete an account and everything it knows: its subscriptions and read state
|
||||
Rm { name: String },
|
||||
/// Rename an account, keeping its feeds, read state and admin rights. This is how an
|
||||
/// account made before the proxy takes the name the proxy signs it in as
|
||||
Rename { name: String, new_name: String },
|
||||
}
|
||||
|
||||
/// What a brand new database starts with, so there is always a way in. Announced loudly
|
||||
@@ -273,11 +276,16 @@ fn user_cmd(ctx: &Arc<Ctx>, cmd: UserCmd) -> Result<()> {
|
||||
println!("no accounts yet: ipx user add <name>");
|
||||
}
|
||||
for u in users {
|
||||
let added = u
|
||||
.created
|
||||
.and_then(|t| chrono::DateTime::from_timestamp(t, 0))
|
||||
.map_or("?".into(), |d| d.format("%Y-%m-%d").to_string());
|
||||
let seen = u.last_login.map_or("never signed in".into(), |t| format!("signed in {}", ago(Some(t))));
|
||||
println!(
|
||||
"{:<20} {:<8} {}",
|
||||
"{:<20} {:<6} {:<11} added {added} {seen}",
|
||||
u.name,
|
||||
if u.is_admin { "admin" } else { "" },
|
||||
if u.pass_hash.is_some() { "password" } else { "proxy only" }
|
||||
if u.pass_hash.is_some() { "password" } else { "proxy only" },
|
||||
);
|
||||
}
|
||||
Ok(())
|
||||
@@ -292,6 +300,22 @@ fn user_cmd(ctx: &Arc<Ctx>, cmd: UserCmd) -> Result<()> {
|
||||
println!("password changed for {name}");
|
||||
Ok(())
|
||||
}
|
||||
UserCmd::Rename { name, new_name } => {
|
||||
let name = name.trim().to_ascii_lowercase();
|
||||
// The same rules as a name the proxy vouches for, or the proxy would never find it.
|
||||
let new_name = crate::auth::name_from_header(&new_name)
|
||||
.ok_or_else(|| anyhow::anyhow!("not a usable name: no commas, semicolons or line breaks"))?;
|
||||
let user = ctx
|
||||
.db
|
||||
.user_by_name(&name)?
|
||||
.ok_or_else(|| anyhow::anyhow!("no such account: {name}"))?;
|
||||
if ctx.db.user_by_name(&new_name)?.is_some() {
|
||||
anyhow::bail!("{new_name} already exists");
|
||||
}
|
||||
ctx.db.rename_user(user.id, &new_name)?;
|
||||
println!("renamed {name} to {new_name}");
|
||||
Ok(())
|
||||
}
|
||||
UserCmd::Rm { name } => {
|
||||
let name = name.trim().to_ascii_lowercase();
|
||||
let user = ctx
|
||||
@@ -315,14 +339,24 @@ async fn run(ctx: &Arc<Ctx>, cmd: Cmd) -> Result<()> {
|
||||
Cmd::Reap { dry_run } => reap(ctx, dry_run, true),
|
||||
Cmd::Download { enclosure } => download_one(ctx, enclosure).await,
|
||||
Cmd::Status => {
|
||||
let (pending, downloaded) = ctx.db.counts()?;
|
||||
let feeds = subscriptions(ctx).map(|s| s.len()).unwrap_or(0);
|
||||
ctx.out.emit(Event::Status { feeds, pending, downloaded });
|
||||
ctx.out.emit(status(ctx));
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The counts `ipx status` prints. A running daemon's socket answers with this directly rather
|
||||
/// than through the job queue.
|
||||
fn status(ctx: &Ctx) -> Event {
|
||||
match ctx.db.counts() {
|
||||
Ok((pending, downloaded)) => {
|
||||
let feeds = subscriptions(ctx).map(|s| s.len()).unwrap_or(0);
|
||||
Event::Status { feeds, pending, downloaded }
|
||||
}
|
||||
Err(e) => Event::Error { msg: format!("{e:#}") },
|
||||
}
|
||||
}
|
||||
|
||||
async fn daemon(
|
||||
ctx: Arc<Ctx>,
|
||||
config_path: PathBuf,
|
||||
@@ -334,8 +368,7 @@ async fn daemon(
|
||||
anyhow::bail!("a daemon is already listening on {}", socket.display());
|
||||
}
|
||||
|
||||
// A database with nobody in it cannot be signed into, and an install that predates
|
||||
// accounts still has to serve its owner. Both get the same starting point.
|
||||
// A database with nobody in it cannot be signed into.
|
||||
if ctx.db.users()?.is_empty() {
|
||||
ctx.db.create_user("admin", Some(&crate::auth::hash_password(DEFAULT_PASSWORD)?), true)?;
|
||||
tracing::warn!(
|
||||
@@ -344,32 +377,51 @@ async fn daemon(
|
||||
);
|
||||
}
|
||||
|
||||
// A library that predates accounts belongs to whoever was using it: the admin.
|
||||
if let Some(admin) = ctx.db.users()?.into_iter().find(|u| u.is_admin) {
|
||||
let catalogue: Vec<String> = ctx.cfg().feeds.keys().cloned().collect();
|
||||
match ctx.db.adopt_existing_library(admin.id, &catalogue) {
|
||||
match ctx.db.adopt_catalogue(admin.id, &catalogue) {
|
||||
Ok(0) => {}
|
||||
Ok(n) => tracing::info!(user = %admin.name, entries = n, "adopted the existing library"),
|
||||
Err(e) => tracing::error!(error = %e, "could not adopt the existing library"),
|
||||
Ok(n) => tracing::info!(user = %admin.name, feeds = n, "subscribed the first admin to the catalogue"),
|
||||
Err(e) => tracing::error!(error = %e, "could not subscribe the first admin to the catalogue"),
|
||||
}
|
||||
}
|
||||
|
||||
match migrate_opml_children(&ctx) {
|
||||
Ok(n) if n > 0 => tracing::info!(count = n, "moved OPML feeds out of config.toml into the database"),
|
||||
Ok(_) => {}
|
||||
Err(e) => tracing::warn!(error = ?e, "could not tidy OPML feeds out of the config"),
|
||||
}
|
||||
|
||||
match ctx.db.requeue_interrupted() {
|
||||
Ok(n) if n > 0 => tracing::info!(count = n, "requeued downloads interrupted by a restart"),
|
||||
Ok(_) => {}
|
||||
Err(e) => tracing::warn!(error = ?e, "could not requeue interrupted downloads"),
|
||||
}
|
||||
|
||||
match retire_stranded(&ctx) {
|
||||
Ok(0) => {}
|
||||
Ok(n) => tracing::info!(feeds = n, "retired feeds whose OPML is no longer in config"),
|
||||
Err(e) => tracing::warn!(error = ?e, "could not retire feeds whose OPML is no longer in config"),
|
||||
}
|
||||
|
||||
// Before the parser knew WordPress's numbered player URLs, a file it listed twice was
|
||||
// downloaded twice. The repeats fold into the first, and their spare copies are deleted.
|
||||
match ctx.db.merge_repeated_enclosures(feed::same_file_key) {
|
||||
Ok((0, _)) => {}
|
||||
Ok((n, spare)) => {
|
||||
for path in &spare {
|
||||
if let Err(e) = std::fs::remove_file(path) {
|
||||
tracing::warn!(path, error = %e, "could not delete a spare copy");
|
||||
}
|
||||
}
|
||||
tracing::info!(enclosures = n, files = spare.len(), "folded files WordPress listed twice");
|
||||
}
|
||||
Err(e) => tracing::warn!(error = ?e, "could not fold files WordPress listed twice"),
|
||||
}
|
||||
|
||||
let (tx_cmd, mut rx_cmd) = mpsc::channel::<Cmd>(64);
|
||||
|
||||
let web = start_web(&ctx, &config_path, web_addr, &tx_cmd, &events).await?;
|
||||
let server = tokio::spawn(ipc::serve(socket.clone(), events.clone(), tx_cmd));
|
||||
// status is answered by the socket itself; everything else waits its turn in the queue.
|
||||
let answer: ipc::StatusFn = {
|
||||
let ctx = ctx.clone();
|
||||
Arc::new(move || status(&ctx))
|
||||
};
|
||||
let server = tokio::spawn(ipc::serve(socket.clone(), events.clone(), tx_cmd, answer));
|
||||
|
||||
// One command at a time: the queue is what keeps two scans from overlapping.
|
||||
let mut ticker = tokio::time::interval(std::time::Duration::from_secs(60));
|
||||
@@ -468,7 +520,7 @@ async fn start_web(
|
||||
let mut fresh = (*cfg).clone();
|
||||
fresh.web.enabled = true;
|
||||
fresh.web.bind = bind.clone();
|
||||
fresh.web.token = web::generate_token();
|
||||
fresh.web.token = crate::auth::new_session_token();
|
||||
fresh.save(config_path)?;
|
||||
ctx.reload_cfg(config_path)?;
|
||||
println!("web ui token generated. Open:\n http://{bind}/?token={}", fresh.web.token);
|
||||
@@ -517,8 +569,9 @@ async fn add(
|
||||
keywords: Vec<String>,
|
||||
) -> Result<()> {
|
||||
let mut cfg = (*ctx.cfg()).clone();
|
||||
let url = &feed::expand_input(url);
|
||||
// Includes feeds derived from an OPML, or the same show could be added twice.
|
||||
if let Some(existing) = subscriptions(ctx)?.iter().find(|s| s.cfg.url == url) {
|
||||
if let Some(existing) = subscriptions(ctx)?.iter().find(|s| feed::same_feed(&s.cfg.url, url)) {
|
||||
anyhow::bail!("already subscribed as {:?}", existing.id);
|
||||
}
|
||||
let id = add_one(ctx, &mut cfg, url, folder, keywords).await?;
|
||||
@@ -549,6 +602,7 @@ pub async fn add_one(
|
||||
username: None,
|
||||
password: None,
|
||||
password_env: None,
|
||||
category: None,
|
||||
};
|
||||
|
||||
let title = match feed::fetch(&ctx.client, &probe, None, None).await {
|
||||
@@ -569,10 +623,17 @@ pub async fn add_one(
|
||||
|
||||
// Slugs must be unique across derived feeds too, or a new feed can collide with one
|
||||
// an OPML already introduced.
|
||||
let taken: std::collections::BTreeMap<String, config::Feed> = subscriptions(ctx)?
|
||||
let mut taken: std::collections::BTreeMap<String, config::Feed> = subscriptions(ctx)?
|
||||
.into_iter()
|
||||
.map(|s| (s.id, s.cfg))
|
||||
.collect();
|
||||
// A removed feed keeps its rows, so its id is only free again for the same feed: re-adding
|
||||
// it gets its history back, and a different feed does not inherit someone else's.
|
||||
for (id, other) in ctx.db.feed_urls()? {
|
||||
if !feed::same_feed(&other, url) {
|
||||
taken.entry(id).or_insert_with(|| probe.clone());
|
||||
}
|
||||
}
|
||||
let id = config::unique_slug(&title, &taken);
|
||||
cfg.feeds.insert(id.clone(), probe);
|
||||
Ok(id)
|
||||
@@ -598,6 +659,7 @@ fn rm(ctx: &Ctx, config_path: &std::path::Path, feed: &str) -> Result<()> {
|
||||
cfg.save(config_path)?;
|
||||
// State and files stay: re-adding the feed should not re-download its back catalogue.
|
||||
println!("removed {feed}; downloads and history kept");
|
||||
retire_group(ctx, feed)?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
@@ -667,6 +729,7 @@ pub fn subscribe_opml(
|
||||
username: None,
|
||||
password: None,
|
||||
password_env: None,
|
||||
category: None,
|
||||
},
|
||||
);
|
||||
grew = true;
|
||||
@@ -803,11 +866,15 @@ async fn fetch(ctx: &Arc<Ctx>, only: Option<&str>, force: bool) -> Result<()> {
|
||||
feed: id.clone(),
|
||||
reason: "not modified".into(),
|
||||
}),
|
||||
Ok(Outcome::Empty) => ctx.out.emit(Event::FeedSkip {
|
||||
feed: id.clone(),
|
||||
reason: "nothing yet".into(),
|
||||
}),
|
||||
Ok(Outcome::Opml { added, removed, kept, total }) => {
|
||||
ctx.out.emit(Event::FeedSkip {
|
||||
feed: id.clone(),
|
||||
reason: format!(
|
||||
"OPML: {total} feed(s) listed, {} added, {removed} unsubscribed, {kept} kept without a listing",
|
||||
"{total} feed(s) listed, {} added, {removed} unsubscribed, {kept} kept without a listing",
|
||||
added.len()
|
||||
),
|
||||
});
|
||||
@@ -880,6 +947,12 @@ pub fn subscriptions(ctx: &Ctx) -> Result<Vec<Sub>> {
|
||||
continue; // promoted to config at some point; that entry wins
|
||||
}
|
||||
let parent = cfg.feeds.get(&m.group_id);
|
||||
if parent.is_none() {
|
||||
// The OPML or Patreon feed this was derived from is no longer in config --
|
||||
// removing it should have retired these rows too (see `retire_group`), but
|
||||
// skip them here regardless so a row that slips through is never scanned.
|
||||
continue;
|
||||
}
|
||||
let base = parent
|
||||
.and_then(|p| p.folder.clone())
|
||||
.or_else(|| ctx.db.feed_summary(&m.group_id).ok().and_then(|s| s.title))
|
||||
@@ -900,6 +973,7 @@ pub fn subscriptions(ctx: &Ctx) -> Result<Vec<Sub>> {
|
||||
username: parent.and_then(|p| p.username.clone()),
|
||||
password: parent.and_then(|p| p.password.clone()),
|
||||
password_env: parent.and_then(|p| p.password_env.clone()),
|
||||
category: None,
|
||||
},
|
||||
managed: true,
|
||||
});
|
||||
@@ -908,34 +982,44 @@ pub fn subscriptions(ctx: &Ctx) -> Result<Vec<Sub>> {
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
/// Moves OPML children that older versions wrote into config.toml over to the database.
|
||||
/// They were never yours to edit, and 80-odd of them made the file unreadable.
|
||||
fn migrate_opml_children(ctx: &Ctx) -> Result<usize> {
|
||||
let cfg = (*ctx.cfg()).clone();
|
||||
let children: Vec<(String, config::Feed)> = cfg
|
||||
.feeds
|
||||
/// Retires every feed derived from `parent_id`, now that nothing subscribes to the OPML or
|
||||
/// Patreon feed that listed them: the same rule `sync_group` applies to one the list drops --
|
||||
/// removed if nothing was downloaded, orphaned and kept otherwise. Called once the parent
|
||||
/// itself is removed, since `subscriptions()` would otherwise keep scanning them under a
|
||||
/// fallback policy meant for a feed with no parent at all. A feed promoted to config is not
|
||||
/// derived any more, so it is only unmanaged.
|
||||
pub fn retire_group(ctx: &Ctx, parent_id: &str) -> Result<()> {
|
||||
let cfg = ctx.cfg();
|
||||
for m in ctx.db.managed_feeds()?.into_iter().filter(|m| m.group_id == parent_id) {
|
||||
if cfg.feeds.contains_key(&m.id) {
|
||||
// Scanned from its config entry and still read. Dropped as derived, its stored
|
||||
// entries would go with it: davewiner's 11 were promoted without being unmanaged.
|
||||
ctx.db.unmanage(&m.id)?;
|
||||
} else if ctx.db.downloaded_count(&m.id).unwrap_or(1) > 0 {
|
||||
ctx.db.set_orphaned(&m.id, true)?;
|
||||
} else {
|
||||
ctx.db.drop_managed(&m.id)?;
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Retires every group whose parent is gone from config, and returns how many derived rows that
|
||||
/// dropped or unmanaged. An OPML removed before `retire_group` existed left its feeds behind:
|
||||
/// davewiner's 922 were skipped by every scan and never cleared, and their stale errors were
|
||||
/// most of the ones stored.
|
||||
fn retire_stranded(ctx: &Ctx) -> Result<usize> {
|
||||
let cfg = ctx.cfg();
|
||||
let before = ctx.db.managed_feeds()?;
|
||||
let stranded: std::collections::BTreeSet<&str> = before
|
||||
.iter()
|
||||
.filter(|(_, f)| f.group.is_some())
|
||||
.map(|(id, f)| (id.clone(), f.clone()))
|
||||
.map(|m| m.group_id.as_str())
|
||||
.filter(|g| !cfg.feeds.contains_key(*g))
|
||||
.collect();
|
||||
if children.is_empty() {
|
||||
return Ok(0);
|
||||
for group in stranded {
|
||||
retire_group(ctx, group)?;
|
||||
}
|
||||
let mut fresh = cfg.clone();
|
||||
for (id, f) in &children {
|
||||
let group = f.group.clone().unwrap_or_default();
|
||||
let title = ctx
|
||||
.db
|
||||
.feed_summary(id)
|
||||
.ok()
|
||||
.and_then(|s| s.title)
|
||||
.unwrap_or_else(|| id.clone());
|
||||
ctx.db.upsert_managed(id, &f.url, &title, &group)?;
|
||||
fresh.feeds.remove(id);
|
||||
}
|
||||
fresh.save(&ctx.config_path)?;
|
||||
ctx.reload_cfg(&ctx.config_path)?;
|
||||
Ok(children.len())
|
||||
Ok(before.len() - ctx.db.managed_feeds()?.len())
|
||||
}
|
||||
|
||||
/// Seconds to wait before re-checking a feed.
|
||||
@@ -962,8 +1046,11 @@ struct Scan {
|
||||
/// What a scan of one feed turned out to be.
|
||||
enum Outcome {
|
||||
NotModified,
|
||||
/// A response with nothing in it -- the British Antarctic Survey answers a 202 with an
|
||||
/// empty body when it has nothing new to publish. Not a parse failure; try again later.
|
||||
Empty,
|
||||
Feed(Scan),
|
||||
/// The URL served an OPML document, so it is a subscription list rather than a feed.
|
||||
/// The URL is a list of feeds rather than a feed: an OPML, or a Patreon creator's shows.
|
||||
Opml { added: Vec<String>, removed: usize, kept: usize, total: usize },
|
||||
}
|
||||
|
||||
@@ -973,6 +1060,31 @@ async fn scan_one(
|
||||
feed_cfg: &config::Feed,
|
||||
state: &db::HttpState,
|
||||
) -> Result<Outcome> {
|
||||
// A Patreon creator with more than one show is a list of feeds, like an OPML.
|
||||
if feed::is_patreon_creator(&feed_cfg.url) {
|
||||
match feed::patreon_shows(&ctx.client, &feed_cfg.url).await {
|
||||
Ok((name, shows)) if shows.len() > 1 => {
|
||||
ctx.db.touch_feed(id, &feed_cfg.url)?;
|
||||
if let Some(name) = name {
|
||||
ctx.db.set_title(id, &name)?;
|
||||
}
|
||||
// Read as one feed before it was split, it listed every show's items in one
|
||||
// heap. The items go; its files and read state move to each show as the show
|
||||
// lists them (`Db::adopt`), so no show comes up empty for want of a URL.
|
||||
ctx.db.clear_entries(id)?;
|
||||
return sync_group(ctx, id, feed_cfg, &shows).await;
|
||||
}
|
||||
Ok(_) => {} // One show: the creator's feed is that show.
|
||||
// Already split: keep the shows it has rather than read the creator as one heap.
|
||||
Err(e) if ctx.db.managed_feeds()?.iter().any(|m| m.group_id == id) => return Err(e),
|
||||
Err(e) => tracing::warn!(
|
||||
feed = id,
|
||||
error = %format!("{e:#}"),
|
||||
"could not list the Patreon shows; reading it as one feed"
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
let mut fetched = feed::fetch(
|
||||
&ctx.client,
|
||||
feed_cfg,
|
||||
@@ -999,6 +1111,11 @@ async fn scan_one(
|
||||
feed::Fetched::Body { bytes, etag, last_modified } => (bytes, etag, last_modified),
|
||||
};
|
||||
|
||||
if bytes.iter().all(u8::is_ascii_whitespace) {
|
||||
ctx.db.touch_feed(id, &feed_cfg.url)?;
|
||||
return Ok(Outcome::Empty);
|
||||
}
|
||||
|
||||
// A subscribed OPML is a list of feeds, not a feed. The original matched on a ".opml"
|
||||
// URL; sniffing the body also catches one served from a URL without that extension.
|
||||
if feed::is_opml(&bytes) {
|
||||
@@ -1015,22 +1132,42 @@ async fn scan_one(
|
||||
last_modified.as_deref(),
|
||||
parsed.ttl_mins,
|
||||
parsed.image.as_deref(),
|
||||
parsed.category.as_deref(),
|
||||
)?;
|
||||
|
||||
let policy = policy_for(ctx, id, feed_cfg)?;
|
||||
if let Some(parent) = &feed_cfg.group {
|
||||
let listed: Vec<(&str, &str)> = parsed
|
||||
.entries
|
||||
.iter()
|
||||
.flat_map(|e| e.enclosures.iter().map(move |x| (e.guid.as_str(), x.url.as_str())))
|
||||
.collect();
|
||||
ctx.db.adopt(parent, id, &listed)?;
|
||||
}
|
||||
// Verdicts are recorded in `state`, so the download queue below is just "everything still
|
||||
// pending". A filter's verdict is looked at again on every scan, though: made once, at
|
||||
// discovery, it outlived the setting behind it, and allowing explicit items afterwards
|
||||
// changed nothing however often the feed was scanned.
|
||||
let skipped = ctx.db.skipped_by_filter(id)?;
|
||||
let mut scan = Scan::default();
|
||||
for entry in &parsed.entries {
|
||||
if ctx.db.record_entry(id, entry)? {
|
||||
scan.new_entries += 1;
|
||||
}
|
||||
for enc in &entry.enclosures {
|
||||
if !ctx.db.record_enclosure(id, &entry.guid, enc)? {
|
||||
continue; // Seen before: downloaded, skipped or deliberately reaped.
|
||||
}
|
||||
// Filters run once, at discovery, and are recorded in `state`. The download
|
||||
// queue below is then just "everything still pending".
|
||||
if let Some(reason) = reject(&ctx.cfg(), feed_cfg, &policy, entry, enc) {
|
||||
ctx.db.mark_enclosure(&enc.url, "skipped", Some(reason))?;
|
||||
let was = if ctx.db.record_enclosure(id, &entry.guid, enc)? {
|
||||
None
|
||||
} else if let Some(reason) = skipped.get(&enc.url) {
|
||||
Some(reason.as_str())
|
||||
} else {
|
||||
continue; // Settled: queued, downloaded, reaped, or another feed's file.
|
||||
};
|
||||
let now = reject(&ctx.cfg(), feed_cfg, &policy, entry, enc);
|
||||
if now != was {
|
||||
match now {
|
||||
Some(reason) => ctx.db.mark_enclosure(&enc.url, "skipped", Some(reason))?,
|
||||
None => ctx.db.mark_enclosure(&enc.url, "pending", None)?,
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1114,12 +1251,6 @@ async fn scan_one(
|
||||
Ok(Outcome::Feed(scan))
|
||||
}
|
||||
|
||||
/// Brings the feed list in step with a subscribed OPML.
|
||||
///
|
||||
/// New entries are added under the OPML's group and folder. An entry that has gone from
|
||||
/// the OPML is unsubscribed *only if nothing was ever downloaded for it* -- otherwise it
|
||||
/// is kept and flagged, because dropping it would orphan files on disk with nothing in
|
||||
/// the UI to explain them.
|
||||
async fn sync_opml(
|
||||
ctx: &Arc<Ctx>,
|
||||
parent_id: &str,
|
||||
@@ -1130,25 +1261,44 @@ async fn sync_opml(
|
||||
if let Some(title) = feed::opml_title(bytes) {
|
||||
ctx.db.set_title(parent_id, &title)?;
|
||||
}
|
||||
sync_group(ctx, parent_id, parent, &listed).await
|
||||
}
|
||||
|
||||
/// Brings the feed list in step with a list of feeds: a subscribed OPML, or a Patreon
|
||||
/// creator's shows.
|
||||
///
|
||||
/// New entries are added under the list's group and folder. An entry that has gone from
|
||||
/// the list is unsubscribed *only if nothing was ever downloaded for it* -- otherwise it
|
||||
/// is kept and flagged, because dropping it would orphan files on disk with nothing in
|
||||
/// the UI to explain them.
|
||||
async fn sync_group(
|
||||
ctx: &Arc<Ctx>,
|
||||
parent_id: &str,
|
||||
parent: &config::Feed,
|
||||
listed: &[(String, String)],
|
||||
) -> Result<Outcome> {
|
||||
let cfg = ctx.cfg();
|
||||
let existing = ctx.db.managed_feeds()?;
|
||||
let mut added = vec![];
|
||||
|
||||
for (title, url) in &listed {
|
||||
for (title, url) in listed {
|
||||
// Already known, whether derived or promoted into the config.
|
||||
if let Some(m) = existing.iter().find(|m| &m.url == url) {
|
||||
ctx.db.upsert_managed(&m.id, url, title, parent_id)?;
|
||||
continue;
|
||||
}
|
||||
if cfg.feeds.values().any(|f| &f.url == url) {
|
||||
// A Patreon show you added by hand may be spelled differently from the one listed.
|
||||
if cfg.feeds.values().any(|f| feed::same_feed(&f.url, url)) {
|
||||
continue;
|
||||
}
|
||||
// A removed feed keeps its rows, so its id is only free again for the same feed.
|
||||
let known = ctx.db.feed_urls()?;
|
||||
let taken: std::collections::BTreeMap<String, config::Feed> = cfg
|
||||
.feeds
|
||||
.keys()
|
||||
.chain(existing.iter().map(|m| &m.id))
|
||||
.chain(added.iter())
|
||||
.chain(known.iter().filter(|(_, u)| !feed::same_feed(u, url)).map(|(id, _)| id))
|
||||
.map(|id| (id.clone(), parent.clone()))
|
||||
.collect();
|
||||
let id = config::unique_slug(title, &taken);
|
||||
@@ -1255,7 +1405,7 @@ pub struct Policy {
|
||||
|
||||
fn policy_for(ctx: &Ctx, id: &str, feed_cfg: &config::Feed) -> Result<Policy> {
|
||||
let global = ctx.cfg().general.max_new_per_check;
|
||||
Ok(merge_policy(&ctx.db.subscribers(id)?, feed_cfg, global))
|
||||
Ok(merge_policy(&ctx.db.subscribers(id, feed_cfg.group.as_deref())?, feed_cfg, global))
|
||||
}
|
||||
|
||||
fn merge_policy(subs: &[db::Sub], feed_cfg: &config::Feed, global: usize) -> Policy {
|
||||
@@ -1526,6 +1676,7 @@ mod tests {
|
||||
username: None,
|
||||
password: None,
|
||||
password_env: None,
|
||||
category: None,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1569,4 +1720,66 @@ mod tests {
|
||||
let p = merge_policy(&[sub(None, Some(false), None), sub(None, Some(true), None)], &feed(), 3);
|
||||
assert!(p.auto_download);
|
||||
}
|
||||
|
||||
fn test_ctx(cfg: config::Config) -> Ctx {
|
||||
Ctx {
|
||||
cfg: std::sync::RwLock::new(Arc::new(cfg)),
|
||||
db: db::Db::memory().unwrap(),
|
||||
client: reqwest::Client::new(),
|
||||
out: Emitter::terminal(),
|
||||
torrents: tokio::sync::OnceCell::new(),
|
||||
torrent_slots: Arc::new(tokio::sync::Semaphore::new(2)),
|
||||
config_path: PathBuf::new(),
|
||||
detach_torrents: false,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_derived_feed_is_not_scanned_once_its_opml_leaves_config() {
|
||||
// davewiner: the OPML subscription left config.toml, but its 922 derived rows
|
||||
// stayed in the database and kept being scanned under the no-parent fallback.
|
||||
let ctx = test_ctx(config::Config::default());
|
||||
ctx.db.upsert_managed("child", "http://x/child.xml", "Child", "gone-opml").unwrap();
|
||||
assert!(
|
||||
subscriptions(&ctx).unwrap().iter().all(|s| s.id != "child"),
|
||||
"a derived feed whose parent is gone from config must not be scanned"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn retiring_a_group_drops_what_was_never_downloaded_and_orphans_the_rest() {
|
||||
let ctx = test_ctx(config::Config::default());
|
||||
ctx.db.upsert_managed("empty", "http://x/empty.xml", "Empty", "parent").unwrap();
|
||||
ctx.db.upsert_managed("has-file", "http://x/has-file.xml", "Has File", "parent").unwrap();
|
||||
let enc = feed::Enclosure { url: "http://x/ep.mp3".into(), mime: None, length: None };
|
||||
ctx.db.record_enclosure("has-file", "g1", &enc).unwrap();
|
||||
ctx.db.mark_downloaded(&enc.url, std::path::Path::new("/downloads/ep.mp3"), 1).unwrap();
|
||||
|
||||
retire_group(&ctx, "parent").unwrap();
|
||||
|
||||
let managed = ctx.db.managed_feeds().unwrap();
|
||||
assert!(!managed.iter().any(|m| m.id == "empty"), "nothing downloaded, so it is forgotten");
|
||||
assert!(managed.iter().any(|m| m.id == "has-file"), "has a file on disk, so it is kept");
|
||||
assert!(ctx.db.feed_summary("has-file").unwrap().orphaned, "and flagged as orphaned");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_stranded_group_is_retired_but_a_promoted_feed_keeps_its_entries() {
|
||||
// davewiner: the OPML left config before retire_group existed, and 11 of its feeds
|
||||
// promoted to config since still said managed = 1.
|
||||
let mut cfg = config::Config::default();
|
||||
cfg.feeds.insert("promoted".into(), feed());
|
||||
cfg.feeds.insert("live-opml".into(), feed());
|
||||
let ctx = test_ctx(cfg);
|
||||
ctx.db.upsert_managed("promoted", "http://x/p.xml", "Promoted", "gone-opml").unwrap();
|
||||
ctx.db.upsert_managed("empty", "http://x/e.xml", "Empty", "gone-opml").unwrap();
|
||||
ctx.db.upsert_managed("listed", "http://x/l.xml", "Listed", "live-opml").unwrap();
|
||||
ctx.db.record_entry("promoted", &feed::Entry { guid: "g1".into(), ..Default::default() }).unwrap();
|
||||
|
||||
assert_eq!(retire_stranded(&ctx).unwrap(), 2, "empty dropped, promoted unmanaged");
|
||||
|
||||
let managed: Vec<String> = ctx.db.managed_feeds().unwrap().into_iter().map(|m| m.id).collect();
|
||||
assert_eq!(managed, ["listed"], "a group still in config is left alone");
|
||||
assert_eq!(ctx.db.feed_summary("promoted").unwrap().entries, 1, "its entries survive");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -154,8 +154,8 @@ mod tests {
|
||||
// One file serves both subscribers, so it takes both of them to release it.
|
||||
let db = Db::memory().unwrap();
|
||||
db.exec_for_test(
|
||||
"INSERT INTO users (id, name, is_admin, created) VALUES (1,'ray',1,0),(2,'sam',0,0);
|
||||
INSERT INTO subscriptions (user_id, feed_id, created) VALUES (1,'f',0),(2,'f',0);
|
||||
"INSERT INTO users (id, name, is_admin) VALUES (1,'ray',1),(2,'sam',0);
|
||||
INSERT INTO subscriptions (user_id, feed_id) VALUES (1,'f'),(2,'f');
|
||||
INSERT INTO entries (feed_id, guid, first_seen) VALUES
|
||||
('f', 'keep', 0),
|
||||
('f', 'half', 0),
|
||||
@@ -189,7 +189,7 @@ mod tests {
|
||||
fn prune_keeps_entries_that_still_have_a_file() {
|
||||
let db = Db::memory().unwrap();
|
||||
db.exec_for_test(
|
||||
"INSERT INTO users (id, name, is_admin, created) VALUES (1,'ray',1,0);
|
||||
"INSERT INTO users (id, name, is_admin) VALUES (1,'ray',1);
|
||||
INSERT INTO entry_state (user_id, feed_id, guid, flagged) VALUES (1,'f','flagged',1);
|
||||
INSERT INTO entries (feed_id, guid, first_seen) VALUES
|
||||
('f', 'has-file', 100),
|
||||
|
||||
297
src/web.rs
297
src/web.rs
@@ -12,9 +12,7 @@ use axum::{
|
||||
},
|
||||
routing::{delete, get, patch, post},
|
||||
};
|
||||
use futures_util::StreamExt;
|
||||
use serde::Deserialize;
|
||||
use tokio_stream::wrappers::BroadcastStream;
|
||||
use tower::ServiceExt;
|
||||
use tower_http::services::ServeFile;
|
||||
use serde::Serialize;
|
||||
@@ -35,28 +33,6 @@ pub struct WebState {
|
||||
pub events: broadcast::Sender<Event>,
|
||||
}
|
||||
|
||||
/// A 32-hex-character shared secret, generated when config.toml has none.
|
||||
///
|
||||
/// ponytail: /dev/urandom rather than a CSPRNG crate -- 16 bytes, once, on a Unix-only
|
||||
/// binary. Falls back to the clock only if urandom is somehow unreadable, which would be a
|
||||
/// weak token, so that case is logged loudly.
|
||||
pub fn generate_token() -> String {
|
||||
use std::io::Read;
|
||||
let mut bytes = [0u8; 16];
|
||||
match std::fs::File::open("/dev/urandom").and_then(|mut f| f.read_exact(&mut bytes)) {
|
||||
Ok(()) => {}
|
||||
Err(e) => {
|
||||
tracing::error!(error = %e, "could not read /dev/urandom; token is NOT secure");
|
||||
let n = std::time::SystemTime::now()
|
||||
.duration_since(std::time::UNIX_EPOCH)
|
||||
.map(|d| d.as_nanos() as u64)
|
||||
.unwrap_or(0);
|
||||
bytes[..8].copy_from_slice(&n.to_le_bytes());
|
||||
}
|
||||
}
|
||||
bytes.iter().map(|b| format!("{b:02x}")).collect()
|
||||
}
|
||||
|
||||
pub fn router(state: WebState) -> Router {
|
||||
Router::new()
|
||||
.route("/", get(index))
|
||||
@@ -88,6 +64,8 @@ pub fn router(state: WebState) -> Router {
|
||||
// Signing in cannot require being signed in, so these sit outside the auth layer.
|
||||
.route("/login", get(login_page))
|
||||
.route("/api/login", post(login))
|
||||
.route("/icon.png", get(icon))
|
||||
.route("/inter.woff2", get(inter))
|
||||
.layer(middleware::from_fn(access_log))
|
||||
.with_state(state)
|
||||
}
|
||||
@@ -114,23 +92,8 @@ async fn auth(State(state): State<WebState>, mut req: Request, next: Next) -> Re
|
||||
let cfg = state.ctx.cfg();
|
||||
let token = cfg.web.token.clone();
|
||||
|
||||
let peer = req
|
||||
.extensions()
|
||||
.get::<axum::extract::ConnectInfo<std::net::SocketAddr>>()
|
||||
.map(|c| c.0.ip().to_string())
|
||||
.unwrap_or_default();
|
||||
|
||||
// 1. A header, but only from a hop we were told to believe. Anyone able to reach the
|
||||
// port could otherwise send it and be whoever they liked.
|
||||
let vouched = (!cfg.web.trusted_header.is_empty()
|
||||
&& cfg.web.trusted_proxies.iter().any(|p| p == &peer))
|
||||
.then(|| {
|
||||
req.headers()
|
||||
.get(&cfg.web.trusted_header)
|
||||
.and_then(|v| v.to_str().ok())
|
||||
.and_then(crate::auth::name_from_header)
|
||||
})
|
||||
.flatten();
|
||||
// 1. A header, but only from a hop we were told to believe.
|
||||
let vouched = vouched_name(&cfg, &req);
|
||||
|
||||
let mut set_cookie: Option<String> = None;
|
||||
let mut user = None;
|
||||
@@ -156,7 +119,13 @@ async fn auth(State(state): State<WebState>, mut req: Request, next: Next) -> Re
|
||||
None
|
||||
}
|
||||
};
|
||||
// Every request comes vouched for; signed_in keeps one an hour. Failing to note the time
|
||||
// must not turn anyone away, so its error goes unanswered.
|
||||
if let Some(u) = &user {
|
||||
let _ = state.ctx.db.signed_in(u.id);
|
||||
}
|
||||
}
|
||||
let by_proxy = user.is_some();
|
||||
|
||||
// 2. A session cookie from signing in here.
|
||||
if user.is_none() {
|
||||
@@ -180,6 +149,10 @@ async fn auth(State(state): State<WebState>, mut req: Request, next: Next) -> Re
|
||||
if supplied.is_some_and(|t| constant_time_eq(&t, &token)) {
|
||||
user = admin_user(&state);
|
||||
if from_query.is_some() {
|
||||
// The token link is a sign-in; the cookie it leaves behind is not one each time.
|
||||
if let Some(u) = &user {
|
||||
let _ = state.ctx.db.signed_in(u.id);
|
||||
}
|
||||
set_cookie = Some(format!(
|
||||
"{COOKIE}={token}; Path=/; HttpOnly; SameSite=Lax; Max-Age=31536000"
|
||||
));
|
||||
@@ -203,6 +176,7 @@ async fn auth(State(state): State<WebState>, mut req: Request, next: Next) -> Re
|
||||
};
|
||||
|
||||
req.extensions_mut().insert(user);
|
||||
req.extensions_mut().insert(Proxied(by_proxy));
|
||||
let mut resp = next.run(req).await;
|
||||
if let Some(c) = set_cookie {
|
||||
if let Ok(v) = header::HeaderValue::from_str(&c) {
|
||||
@@ -214,6 +188,29 @@ async fn auth(State(state): State<WebState>, mut req: Request, next: Next) -> Re
|
||||
|
||||
const SESSION_COOKIE: &str = "ipx_session";
|
||||
|
||||
/// Whether the proxy signed this request in, rather than a session or the token: signing out
|
||||
/// has to go through the proxy then, or its next request signs the person straight back in.
|
||||
#[derive(Clone, Copy)]
|
||||
struct Proxied(bool);
|
||||
|
||||
/// The name the proxy vouches for, when this request came from one of `trusted_proxies` and
|
||||
/// carries `trusted_header`. Anyone able to reach the port could otherwise send the header and
|
||||
/// be whoever they liked.
|
||||
fn vouched_name(cfg: &crate::config::Config, req: &Request) -> Option<String> {
|
||||
let peer = req
|
||||
.extensions()
|
||||
.get::<axum::extract::ConnectInfo<std::net::SocketAddr>>()
|
||||
.map(|c| c.0.ip().to_string())
|
||||
.unwrap_or_default();
|
||||
if cfg.web.trusted_header.is_empty() || !cfg.web.trusted_proxies.iter().any(|p| p == &peer) {
|
||||
return None;
|
||||
}
|
||||
req.headers()
|
||||
.get(&cfg.web.trusted_header)
|
||||
.and_then(|v| v.to_str().ok())
|
||||
.and_then(crate::auth::name_from_header)
|
||||
}
|
||||
|
||||
/// Handlers take `User` to say they need one; the auth layer put it there, and nothing
|
||||
/// reaches a handler without passing through it.
|
||||
impl<S: Send + Sync> axum::extract::FromRequestParts<S> for crate::db::User {
|
||||
@@ -276,6 +273,7 @@ async fn login(
|
||||
let user = user.expect("verified above");
|
||||
let token = crate::auth::new_session_token();
|
||||
state.ctx.db.create_session(user.id, &token)?;
|
||||
state.ctx.db.signed_in(user.id)?;
|
||||
tracing::info!(user = %user.name, "signed in");
|
||||
|
||||
let days = state.ctx.cfg().web.session_days.max(1);
|
||||
@@ -306,8 +304,15 @@ async fn logout(State(state): State<WebState>, req: Request) -> Response {
|
||||
resp
|
||||
}
|
||||
|
||||
async fn me(user: crate::db::User) -> Json<serde_json::Value> {
|
||||
Json(serde_json::json!({ "name": user.name, "admin": user.is_admin }))
|
||||
/// Who is signed in, and, for someone the proxy signed in, where Sign out should send them.
|
||||
async fn me(
|
||||
State(state): State<WebState>,
|
||||
user: crate::db::User,
|
||||
axum::Extension(Proxied(by_proxy)): axum::Extension<Proxied>,
|
||||
) -> Json<serde_json::Value> {
|
||||
let url = state.ctx.cfg().web.sign_out_url.clone();
|
||||
let sign_out = (by_proxy && !url.is_empty()).then_some(url);
|
||||
Json(serde_json::json!({ "name": user.name, "admin": user.is_admin, "sign_out": sign_out }))
|
||||
}
|
||||
|
||||
// ---- accounts: admin only ----
|
||||
@@ -340,6 +345,7 @@ async fn list_users(
|
||||
.map(|u| {
|
||||
serde_json::json!({
|
||||
"id": u.id, "name": u.name, "admin": u.is_admin, "password": u.pass_hash.is_some(),
|
||||
"created": u.created, "last_login": u.last_login,
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
@@ -429,8 +435,32 @@ async fn remove_user(
|
||||
Ok(StatusCode::NO_CONTENT)
|
||||
}
|
||||
|
||||
async fn login_page() -> Html<&'static str> {
|
||||
Html(include_str!("../web/login.html"))
|
||||
/// The password form, except for someone the proxy vouches for: they are signed in already, and
|
||||
/// the form only made it look as if they were not.
|
||||
async fn login_page(State(state): State<WebState>, req: Request) -> Response {
|
||||
if vouched_name(&state.ctx.cfg(), &req).is_some() {
|
||||
return Redirect::to("/").into_response();
|
||||
}
|
||||
Html(include_str!("../web/login.html")).into_response()
|
||||
}
|
||||
|
||||
/// The 2004 icon, served once for both pages rather than inlined as base64 into each. The
|
||||
/// sign-in page shows it, so it sits outside the auth layer with /login.
|
||||
async fn icon() -> impl IntoResponse {
|
||||
(
|
||||
[(header::CONTENT_TYPE, "image/png"), (header::CACHE_CONTROL, "max-age=86400")],
|
||||
include_bytes!("../web/ipodderx-icon.png").as_slice(),
|
||||
)
|
||||
}
|
||||
|
||||
/// Inter, the pages' typeface, served from the binary as the icon is, so neither page loads
|
||||
/// anything from anyone else. Outside the auth layer for the sign-in page. Its licence, the SIL
|
||||
/// Open Font License, is web/Inter-LICENSE.txt.
|
||||
async fn inter() -> impl IntoResponse {
|
||||
(
|
||||
[(header::CONTENT_TYPE, "font/woff2"), (header::CACHE_CONTROL, "max-age=604800")],
|
||||
include_bytes!("../web/InterVariable.woff2").as_slice(),
|
||||
)
|
||||
}
|
||||
|
||||
fn constant_time_eq(a: &str, b: &str) -> bool {
|
||||
@@ -452,6 +482,9 @@ struct FeedRow {
|
||||
title: Option<String>,
|
||||
image: Option<String>,
|
||||
folder: Option<String>,
|
||||
/// The Directory category an admin gave it, and the one the feed names itself, which wins.
|
||||
category: Option<String>,
|
||||
feed_category: Option<String>,
|
||||
keywords: Vec<String>,
|
||||
allow_explicit: bool,
|
||||
auto_download: bool,
|
||||
@@ -469,6 +502,9 @@ struct FeedRow {
|
||||
last_checked: Option<i64>,
|
||||
next_check: Option<i64>,
|
||||
last_error: Option<String>,
|
||||
/// Set once `last_error` is a kind worth telling someone about and it has held for a
|
||||
/// day -- a feed that fails once and reads fine an hour later (macmanx) never gets here.
|
||||
failing: Option<FailingRow>,
|
||||
entries: i64,
|
||||
downloaded: i64,
|
||||
unread: i64,
|
||||
@@ -476,6 +512,15 @@ struct FeedRow {
|
||||
subscribers: i64,
|
||||
}
|
||||
|
||||
#[derive(Serialize)]
|
||||
struct FailingRow {
|
||||
reason: &'static str,
|
||||
new_url: Option<String>,
|
||||
}
|
||||
|
||||
/// A day, in seconds: how long an error has to hold before the UI mentions it.
|
||||
const FLAG_AFTER_SECS: i64 = 86_400;
|
||||
|
||||
async fn feeds(
|
||||
State(state): State<WebState>,
|
||||
user: crate::db::User,
|
||||
@@ -495,6 +540,9 @@ async fn feeds(
|
||||
let mut out = Vec::with_capacity(mine.len());
|
||||
for sub in &subs {
|
||||
let (id, feed) = (&sub.id, &sub.cfg);
|
||||
// In a group, what you have not set on the feed comes from your settings on the group,
|
||||
// the same fallback the scanner uses (`Db::subscribers`).
|
||||
let up = feed.group.as_deref().and_then(|g| mine.get(g));
|
||||
let Some(mine) = mine.get(id) else { continue };
|
||||
let s = state.ctx.db.feed_summary(id)?;
|
||||
let st = state.ctx.db.http_state(id)?;
|
||||
@@ -504,11 +552,24 @@ async fn feeds(
|
||||
title: s.title,
|
||||
image: s.image,
|
||||
folder: feed.folder.clone(),
|
||||
keywords: mine.keywords.clone().unwrap_or_else(|| feed.keywords.clone()),
|
||||
allow_explicit: mine.allow_explicit.unwrap_or(feed.allow_explicit),
|
||||
auto_download: mine.auto_download.unwrap_or(feed.auto_download),
|
||||
category: feed.category.clone(),
|
||||
feed_category: s.category,
|
||||
keywords: mine
|
||||
.keywords
|
||||
.clone()
|
||||
.or_else(|| up.and_then(|u| u.keywords.clone()))
|
||||
.unwrap_or_else(|| feed.keywords.clone()),
|
||||
allow_explicit: mine
|
||||
.allow_explicit
|
||||
.or(up.and_then(|u| u.allow_explicit))
|
||||
.unwrap_or(feed.allow_explicit),
|
||||
auto_download: mine
|
||||
.auto_download
|
||||
.or(up.and_then(|u| u.auto_download))
|
||||
.unwrap_or(feed.auto_download),
|
||||
max_new_per_check: mine
|
||||
.max_new_per_check
|
||||
.or(up.and_then(|u| u.max_new_per_check))
|
||||
.map(|n| n as usize)
|
||||
.or(feed.max_new_per_check),
|
||||
group: feed.group.clone(),
|
||||
@@ -525,6 +586,12 @@ async fn feeds(
|
||||
next_check: s
|
||||
.last_checked
|
||||
.map(|t| t + crate::due_after(&cfg, feed, st.ttl_mins) as i64),
|
||||
failing: s
|
||||
.error_since
|
||||
.filter(|since| crate::db::now() - since >= FLAG_AFTER_SECS)
|
||||
.and_then(|_| s.last_error.as_deref())
|
||||
.and_then(crate::feed::explain_failure)
|
||||
.map(|f| FailingRow { reason: f.reason, new_url: f.new_url }),
|
||||
last_error: s.last_error,
|
||||
entries: s.entries,
|
||||
downloaded: s.downloaded,
|
||||
@@ -576,28 +643,52 @@ struct PopularRow {
|
||||
subscribers: i64,
|
||||
/// Yours already. Everyone counts, you included, so your own feeds are listed too.
|
||||
subscribed: bool,
|
||||
/// The feed's own iTunes category, if it names one; most blogs do not.
|
||||
category: Option<String>,
|
||||
/// Any audio or video enclosure. Unlike category, every feed has an answer, so the
|
||||
/// Directory's Podcasts and Blogs between them hold everything.
|
||||
podcast: bool,
|
||||
}
|
||||
|
||||
/// Every feed that may be listed, with everyone counted, you included, most subscribers
|
||||
/// first. Popular is the top of it, the directory is all of it, and it is all that
|
||||
/// `subscribe_popular` will subscribe you to.
|
||||
/// `subscribe_popular` will subscribe you to. An OPML or a Patreon creator is listed as the
|
||||
/// feeds inside it and never itself: both lists are for finding a show.
|
||||
fn popular(state: &WebState, user_id: i64) -> Result<Vec<PopularRow>> {
|
||||
let db = &state.ctx.db;
|
||||
let mine: std::collections::HashSet<String> =
|
||||
db.subscriptions_for(user_id)?.into_iter().map(|s| s.feed_id).collect();
|
||||
let counts = db.subscriber_counts()?;
|
||||
let media = db.media_feeds()?;
|
||||
let catalogue = crate::subscriptions(&state.ctx)?;
|
||||
let by_id: std::collections::HashMap<&str, &crate::config::Feed> =
|
||||
catalogue.iter().map(|s| (s.id.as_str(), &s.cfg)).collect();
|
||||
let is_folder: std::collections::HashSet<&str> =
|
||||
catalogue.iter().filter_map(|s| s.cfg.group.as_deref()).collect();
|
||||
let mut out = vec![];
|
||||
for s in crate::subscriptions(&state.ctx)? {
|
||||
for s in &catalogue {
|
||||
let n = counts.get(&s.id).copied().unwrap_or(0);
|
||||
// A feed from an OPML rides on the OPML: everyone subscribed to it counts every feed
|
||||
// inside, which would bury everything anyone chose on purpose.
|
||||
let from_opml = s.managed || s.cfg.group.is_some();
|
||||
if n == 0 || from_opml || looks_private(&s.cfg) {
|
||||
// A feed inside an OPML that looks private is as private as the OPML.
|
||||
let folder = s.cfg.group.as_deref().and_then(|g| by_id.get(g));
|
||||
if n == 0
|
||||
|| is_folder.contains(s.id.as_str())
|
||||
|| looks_private(&s.cfg)
|
||||
|| folder.is_some_and(|f| looks_private(f))
|
||||
{
|
||||
continue;
|
||||
}
|
||||
let sum = db.feed_summary(&s.id)?;
|
||||
let subscribed = mine.contains(&s.id);
|
||||
out.push(PopularRow { id: s.id, title: sum.title, image: sum.image, subscribers: n, subscribed });
|
||||
out.push(PopularRow {
|
||||
id: s.id.clone(),
|
||||
title: sum.title,
|
||||
image: sum.image,
|
||||
subscribers: n,
|
||||
subscribed,
|
||||
// The feed's own wins; an admin's is for the feeds, mostly blogs, that name none.
|
||||
category: sum.category.or_else(|| s.cfg.category.clone()),
|
||||
podcast: media.contains(&s.id),
|
||||
});
|
||||
}
|
||||
out.sort_by(|a, b| b.subscribers.cmp(&a.subscribers).then_with(|| sort_name(a).cmp(&sort_name(b))));
|
||||
Ok(out)
|
||||
@@ -612,7 +703,7 @@ async fn get_popular(
|
||||
Ok(Json(rows))
|
||||
}
|
||||
|
||||
/// Every feed that may be listed, A to Z.
|
||||
/// Every feed that may be listed, A to Z, with the feeds inside an OPML in place of the OPML.
|
||||
async fn get_directory(
|
||||
State(state): State<WebState>,
|
||||
user: crate::db::User,
|
||||
@@ -634,7 +725,7 @@ async fn subscribe_popular(
|
||||
Path(id): Path<String>,
|
||||
) -> Result<Json<serde_json::Value>, ApiError> {
|
||||
if !popular(&state, user.id)?.iter().any(|p| p.id == id) {
|
||||
return Err(ApiError::bad_request(format!("{id:?} is not on the popular list")));
|
||||
return Err(ApiError::bad_request(format!("{id:?} is not in the directory")));
|
||||
}
|
||||
state.ctx.db.subscribe(user.id, &id)?;
|
||||
Ok(Json(serde_json::json!({ "id": id })))
|
||||
@@ -719,6 +810,7 @@ mod tests {
|
||||
username: None,
|
||||
password: None,
|
||||
password_env: None,
|
||||
category: None,
|
||||
};
|
||||
assert!(!looks_private(&f("https://feeds.twit.tv/twit.xml")));
|
||||
assert!(!looks_private(&f("https://example.com/rss?format=mp3")));
|
||||
@@ -739,7 +831,14 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn only_the_last_admin_is_protected() {
|
||||
let u = |id, is_admin| crate::db::User { id, name: format!("u{id}"), pass_hash: None, is_admin };
|
||||
let u = |id, is_admin| crate::db::User {
|
||||
id,
|
||||
name: format!("u{id}"),
|
||||
pass_hash: None,
|
||||
is_admin,
|
||||
created: None,
|
||||
last_login: None,
|
||||
};
|
||||
assert!(last_admin(&[u(1, true), u(2, false)], 1));
|
||||
assert!(!last_admin(&[u(1, true), u(2, true)], 1), "another admin remains");
|
||||
assert!(!last_admin(&[u(1, true), u(2, false)], 2), "not an admin at all");
|
||||
@@ -758,7 +857,7 @@ mod tests {
|
||||
crate::config::Feed {
|
||||
url: url.into(), folder: None, group: None, media_types: None, schedule: None, keywords: vec![], allow_explicit: false,
|
||||
auto_download: true, max_new_per_check: None, username: None,
|
||||
password: None, password_env: None,
|
||||
password: None, password_env: None, category: None,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -770,9 +869,10 @@ mod tests {
|
||||
assert!(absent.schedule.is_none() && absent.folder.is_none());
|
||||
|
||||
let cleared: FeedPatch =
|
||||
serde_json::from_str(r#"{"schedule":null,"folder":null,"max_new_per_check":null}"#).unwrap();
|
||||
serde_json::from_str(r#"{"schedule":null,"folder":null,"category":null,"max_new_per_check":null}"#).unwrap();
|
||||
assert_eq!(cleared.schedule, Some(None), "null must mean clear");
|
||||
assert_eq!(cleared.folder, Some(None));
|
||||
assert_eq!(cleared.category, Some(None));
|
||||
assert_eq!(cleared.max_new_per_check, Some(None));
|
||||
|
||||
let set: FeedPatch = serde_json::from_str(r#"{"schedule":"every 6h"}"#).unwrap();
|
||||
@@ -803,15 +903,6 @@ mod tests {
|
||||
"another feed already has that URL"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn generated_tokens_are_32_hex_chars_and_not_repeated() {
|
||||
let a = generate_token();
|
||||
let b = generate_token();
|
||||
assert_eq!(a.len(), 32);
|
||||
assert!(a.chars().all(|c| c.is_ascii_hexdigit()));
|
||||
assert_ne!(a, b);
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Deserialize)]
|
||||
@@ -877,9 +968,13 @@ fn entry_page(
|
||||
let mut rows =
|
||||
db.entries_in(user_id, feed, filter, search, page.offset, page.limit.clamp(1, 200), &order)?;
|
||||
// Feed HTML is untrusted: it reaches the page only after ammonia has been through it.
|
||||
// Every link opens in a new tab -- ammonia's default rel="noopener noreferrer" already
|
||||
// keeps that safe -- so following one in show notes never navigates away from ipx.
|
||||
let mut sanitizer = ammonia::Builder::new();
|
||||
sanitizer.add_tag_attributes("a", &["target"]).set_tag_attribute_value("a", "target", "_blank");
|
||||
for row in &mut rows {
|
||||
if let Some(d) = &row.description {
|
||||
row.description = Some(ammonia::clean(d));
|
||||
row.description = Some(sanitizer.clean(d).to_string());
|
||||
}
|
||||
}
|
||||
let total = db.count_in(user_id, feed, filter, search)?;
|
||||
@@ -893,6 +988,19 @@ struct NewFeed {
|
||||
folder: Option<String>,
|
||||
#[serde(default)]
|
||||
keywords: Vec<String>,
|
||||
#[serde(default)]
|
||||
allow_explicit: bool,
|
||||
}
|
||||
|
||||
/// The Add feed dialog's explicit box. Like everything on a feed's own dialog it is yours, so it
|
||||
/// goes on your subscription, and before the first scan, which would otherwise skip every
|
||||
/// explicit item.
|
||||
fn explicit_on_add(state: &WebState, user_id: i64, feed_id: &str, allow: bool) -> Result<(), ApiError> {
|
||||
if allow {
|
||||
let sub = crate::db::Sub { feed_id: feed_id.to_owned(), allow_explicit: Some(true), ..Default::default() };
|
||||
state.ctx.db.set_subscription(user_id, &sub)?;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
async fn add_feed(
|
||||
@@ -901,23 +1009,28 @@ async fn add_feed(
|
||||
Json(body): Json<NewFeed>,
|
||||
) -> Result<Json<serde_json::Value>, ApiError> {
|
||||
let mut cfg = (*state.ctx.cfg()).clone();
|
||||
let url = crate::feed::expand_input(&body.url);
|
||||
// Someone else may already have it. Then adding costs nothing: no second fetch, no
|
||||
// second copy on disk, just another name against the same feed.
|
||||
if let Some(existing) = crate::subscriptions(&state.ctx)?
|
||||
.into_iter()
|
||||
.find(|s| s.cfg.url == body.url)
|
||||
.find(|s| crate::feed::same_feed(&s.cfg.url, &url))
|
||||
{
|
||||
let already = state.ctx.db.subscription(user.id, &existing.id)?.is_some();
|
||||
state.ctx.db.subscribe(user.id, &existing.id)?;
|
||||
if !already {
|
||||
explicit_on_add(&state, user.id, &existing.id, body.allow_explicit)?;
|
||||
}
|
||||
scan_soon(&state, Some(existing.id.clone())).await;
|
||||
return Ok(Json(
|
||||
serde_json::json!({ "id": existing.id, "existing": already }),
|
||||
));
|
||||
}
|
||||
let id = crate::add_one(&state.ctx, &mut cfg, &body.url, body.folder, body.keywords).await?;
|
||||
let id = crate::add_one(&state.ctx, &mut cfg, &url, body.folder, body.keywords).await?;
|
||||
cfg.save(&state.config_path)?;
|
||||
state.ctx.reload_cfg(&state.config_path)?;
|
||||
state.ctx.db.subscribe(user.id, &id)?;
|
||||
explicit_on_add(&state, user.id, &id, body.allow_explicit)?;
|
||||
scan_soon(&state, Some(id.clone())).await;
|
||||
Ok(Json(serde_json::json!({ "id": id, "existing": false })))
|
||||
}
|
||||
@@ -944,6 +1057,8 @@ struct FeedPatch {
|
||||
schedule: Option<Option<String>>,
|
||||
#[serde(default, deserialize_with = "double_option")]
|
||||
folder: Option<Option<String>>,
|
||||
#[serde(default, deserialize_with = "double_option")]
|
||||
category: Option<Option<String>>,
|
||||
keywords: Option<Vec<String>>,
|
||||
allow_explicit: Option<bool>,
|
||||
auto_download: Option<bool>,
|
||||
@@ -997,13 +1112,14 @@ async fn patch_feed(
|
||||
|
||||
// The rest describes the feed itself -- where its files land, its address, when it is
|
||||
// polled -- and there is one of those however many people read it.
|
||||
let feed_level = body.url.is_some() || body.folder.is_some() || body.schedule.is_some();
|
||||
let feed_level =
|
||||
body.url.is_some() || body.folder.is_some() || body.schedule.is_some() || body.category.is_some();
|
||||
if !feed_level {
|
||||
return Ok(StatusCode::NO_CONTENT);
|
||||
}
|
||||
if !user.is_admin {
|
||||
return Err(ApiError::forbidden(
|
||||
"the feed's address, folder and schedule are the same for everyone, so only an admin changes them",
|
||||
"the feed's address, folder, schedule and category are the same for everyone, so only an admin changes them",
|
||||
));
|
||||
}
|
||||
let mut cfg = (*state.ctx.cfg()).clone();
|
||||
@@ -1050,6 +1166,9 @@ async fn patch_feed(
|
||||
if let Some(v) = body.folder {
|
||||
feed.folder = v.filter(|s| !s.trim().is_empty());
|
||||
}
|
||||
if let Some(v) = body.category {
|
||||
feed.category = v.map(|s| s.trim().to_owned()).filter(|s| !s.is_empty());
|
||||
}
|
||||
cfg.save(&state.config_path)?;
|
||||
state.ctx.reload_cfg(&state.config_path)?;
|
||||
if url_changed {
|
||||
@@ -1074,7 +1193,7 @@ async fn remove_feed(
|
||||
{
|
||||
state.ctx.db.unsubscribe(user.id, &child.id)?;
|
||||
}
|
||||
if state.ctx.db.subscriber_count(&id)? > 0 {
|
||||
if state.ctx.db.subscriber_counts()?.contains_key(&id) {
|
||||
return Ok(StatusCode::NO_CONTENT);
|
||||
}
|
||||
|
||||
@@ -1089,6 +1208,7 @@ async fn remove_feed(
|
||||
}
|
||||
cfg.save(&state.config_path)?;
|
||||
state.ctx.reload_cfg(&state.config_path)?;
|
||||
crate::retire_group(&state.ctx, &id)?;
|
||||
Ok(StatusCode::NO_CONTENT)
|
||||
}
|
||||
|
||||
@@ -1164,9 +1284,9 @@ async fn delete_file(
|
||||
let complaint = match (starred, unread) {
|
||||
(0, 0) => None,
|
||||
(0, u) => Some(format!("{} subscribed to this feed {} not played it yet", people(u), if u == 1 { "has" } else { "have" })),
|
||||
(st, 0) => Some(format!("another {} starred it to keep", people(st))),
|
||||
(st, 0) => Some(format!("another {} pinned it", people(st))),
|
||||
(st, u) => Some(format!(
|
||||
"another {} starred it to keep, and {} not played it yet",
|
||||
"another {} pinned it, and {} not played it yet",
|
||||
people(st),
|
||||
if u == 1 { "one person has".to_string() } else { format!("{u} have") }
|
||||
)),
|
||||
@@ -1215,9 +1335,20 @@ async fn fetch_now(
|
||||
|
||||
/// The same broadcast the socket clients read, as server-sent events.
|
||||
async fn events(State(state): State<WebState>) -> Sse<impl futures_util::Stream<Item = Result<SseEvent, std::convert::Infallible>>> {
|
||||
let stream = BroadcastStream::new(state.events.subscribe()).filter_map(|ev| async move {
|
||||
let ev = ev.ok()?;
|
||||
Some(Ok(SseEvent::default().data(serde_json::to_string(&ev).ok()?)))
|
||||
// A client that falls behind skips what it missed rather than being cut off.
|
||||
let stream = futures_util::stream::unfold(state.events.subscribe(), |mut rx| async move {
|
||||
loop {
|
||||
match rx.recv().await {
|
||||
Ok(ev) => {
|
||||
if let Ok(data) = serde_json::to_string(&ev) {
|
||||
let ev = Ok::<_, std::convert::Infallible>(SseEvent::default().data(data));
|
||||
return Some((ev, rx));
|
||||
}
|
||||
}
|
||||
Err(broadcast::error::RecvError::Lagged(_)) => {}
|
||||
Err(broadcast::error::RecvError::Closed) => return None,
|
||||
}
|
||||
}
|
||||
});
|
||||
Sse::new(stream).keep_alive(axum::response::sse::KeepAlive::default())
|
||||
}
|
||||
@@ -1243,6 +1374,8 @@ async fn media(
|
||||
#[derive(Deserialize)]
|
||||
struct Position {
|
||||
secs: i64,
|
||||
/// The length the player measured, for an episode whose feed gives none.
|
||||
duration: Option<i64>,
|
||||
}
|
||||
|
||||
async fn set_position(
|
||||
@@ -1251,7 +1384,7 @@ async fn set_position(
|
||||
user: crate::db::User,
|
||||
Json(body): Json<Position>,
|
||||
) -> Result<StatusCode, ApiError> {
|
||||
state.ctx.db.set_position(user.id, &feed_id, &guid, body.secs)?;
|
||||
state.ctx.db.set_position(user.id, &feed_id, &guid, body.secs, body.duration)?;
|
||||
Ok(StatusCode::NO_CONTENT)
|
||||
}
|
||||
|
||||
@@ -1428,8 +1561,6 @@ async fn patch_settings(
|
||||
)));
|
||||
}
|
||||
cfg.general.schedule = sched;
|
||||
// The legacy key would otherwise keep shadowing intent in the file.
|
||||
cfg.general.interval_mins = None;
|
||||
}
|
||||
if let Some(v) = body.max_new_per_check {
|
||||
cfg.general.max_new_per_check = v;
|
||||
|
||||
@@ -6,6 +6,8 @@
|
||||
<description>A synthetic feed used by the parser tests.</description>
|
||||
<ttl>45</ttl>
|
||||
<itunes:explicit>no</itunes:explicit>
|
||||
<itunes:category text="Technology"><itunes:category text="Podcasting"/></itunes:category>
|
||||
<itunes:category text="News"/>
|
||||
|
||||
<item>
|
||||
<title>Episode One</title>
|
||||
|
||||
@@ -88,8 +88,10 @@ const drive = [
|
||||
['opmlModal', () => ctx.opmlModal()],
|
||||
['selectFeed (directory)', () => ctx.selectFeed(':directory')],
|
||||
['selectFeed (popular)', () => ctx.selectFeed(':popular')],
|
||||
['selectFeed (currently listening)', () => ctx.selectFeed(':listening')],
|
||||
['selectFeed (all subscriptions)', () => ctx.selectFeed(':all')],
|
||||
['logsModal', () => ctx.logsModal()],
|
||||
['keysModal', () => ctx.keysModal()],
|
||||
// `const S` is not reachable from here: top-level const/let do not become properties
|
||||
// of a vm context the way var and function declarations do.
|
||||
['renderGroup', () => ctx.renderGroup(feed, [{ ...feed, id: 'child', group: 'f', orphaned: true }])],
|
||||
|
||||
@@ -12,7 +12,7 @@ test('the page loads and lists the configured feeds', async ({ page }) => {
|
||||
// empty, with every handler below the error dead. Server-side checks all passed.
|
||||
// Four top-level feeds in the fixture config; the OPML's children are inside a closed folder.
|
||||
await expect(page.locator('.feed')).toHaveCount(5, { timeout: 15_000 });
|
||||
await expect(page.getByText('Test Show')).toBeVisible();
|
||||
await expect(page.locator('.feed', { hasText: 'Test Show' })).toBeVisible();
|
||||
const errors = [];
|
||||
page.on('pageerror', e => errors.push(e.message));
|
||||
await page.reload();
|
||||
@@ -33,7 +33,7 @@ test('the theme button steps through dark, light and classic, and remembers', as
|
||||
const theme = () => page.evaluate(() => document.documentElement.dataset.theme);
|
||||
for (let i = 0; i < 3 && (await theme()) !== 'classic'; i++) await page.locator('#theme').click();
|
||||
expect(await theme()).toBe('classic');
|
||||
await expect(page.locator('#theme')).toHaveAttribute('title', /Classic.*Click for Dark/);
|
||||
await expect(page.locator('#theme')).toHaveAttribute('title', /Classic.*Click for Auto/);
|
||||
|
||||
await page.reload();
|
||||
await expect.poll(theme).toBe('classic');
|
||||
@@ -41,6 +41,28 @@ test('the theme button steps through dark, light and classic, and remembers', as
|
||||
expect(await page.evaluate(() => getComputedStyle(document.body).fontFamily)).toContain('Lucida Grande');
|
||||
});
|
||||
|
||||
test('the theme dropdown in Settings jumps straight to a theme, including Auto', async ({ page }) => {
|
||||
const theme = () => page.evaluate(() => document.documentElement.dataset.theme);
|
||||
await page.locator('#prefs').click();
|
||||
await expect(page.locator('#stheme')).toHaveValue(await theme());
|
||||
|
||||
await page.locator('#stheme').selectOption('auto');
|
||||
await expect.poll(theme).toBe('auto');
|
||||
// Auto follows the system; emulating a light system must show the light palette live,
|
||||
// no reload needed, since it is a media query rather than something JS picks per click.
|
||||
await page.emulateMedia({ colorScheme: 'light' });
|
||||
await expect.poll(() => page.evaluate(() => getComputedStyle(document.body).backgroundColor))
|
||||
.toBe('rgb(242, 244, 247)'); // --bg in the light palette
|
||||
await page.emulateMedia({ colorScheme: 'dark' });
|
||||
await expect.poll(() => page.evaluate(() => getComputedStyle(document.body).backgroundColor))
|
||||
.toBe('rgb(14, 19, 27)'); // the bare :root is already dark; Auto adds nothing here
|
||||
|
||||
// The header button and the dropdown are the same one setting, not two.
|
||||
await page.locator('#modalCard .cardacts .btn').first().click(); // Cancel, closing the modal
|
||||
await page.locator('#theme').click();
|
||||
expect(await theme()).toBe('dark');
|
||||
});
|
||||
|
||||
test('settings opens and saves the global schedule', async ({ page }) => {
|
||||
await page.locator('#prefs').click();
|
||||
await expect(page.locator('#modal.on')).toBeVisible();
|
||||
@@ -57,7 +79,7 @@ test('settings opens and saves the global schedule', async ({ page }) => {
|
||||
});
|
||||
|
||||
test('episodes show with their metadata, and the text opens below', async ({ page }) => {
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
await expect(page.locator('.ep').first()).toBeVisible({ timeout: 20_000 });
|
||||
await expect(page.getByText('First Episode')).toBeVisible();
|
||||
// Newest first, so target the episode by name rather than by position.
|
||||
@@ -74,7 +96,7 @@ test('episodes show with their metadata, and the text opens below', async ({ pag
|
||||
});
|
||||
|
||||
test('the three panes are there and the item text lands in the bottom one', async ({ page }) => {
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
await expect(page.locator('#list')).toBeVisible();
|
||||
await expect(page.locator('#grab')).toBeVisible(); // the draggable divider
|
||||
await expect(page.locator('#detail')).toContainText('Pick an item');
|
||||
@@ -133,8 +155,53 @@ test('an item with several enclosures lists them all', async ({ page }) => {
|
||||
await expect(page.locator('#files .encbox').nth(1).locator('.kind[title^="image"]')).toBeVisible();
|
||||
});
|
||||
|
||||
test('Currently Listening, its own place below Popular, resumes an episode you started or forgets it', async ({ page }) => {
|
||||
// Second Episode (900s) is 42 seconds in and unfinished. An earlier test may have opened it,
|
||||
// and opening marks it read; it is listed all the same, because read is not finished. This
|
||||
// test used to set it unread first, which hid exactly the bug in issue #14.
|
||||
await page.evaluate(() =>
|
||||
api('/api/entries/test-show/ui-2/flags', { method: 'POST', body: JSON.stringify({ read: true }) }));
|
||||
await page.evaluate(() =>
|
||||
api('/api/entries/test-show/ui-2/position', { method: 'POST', body: JSON.stringify({ secs: 42 }) }));
|
||||
|
||||
// Popular lists feeds and nothing else; the episodes have a place of their own under it.
|
||||
await page.locator('#feedlist .place', { hasText: 'Popular' }).click();
|
||||
await expect(page.locator('#popular')).toBeVisible();
|
||||
await expect(page.locator('#listening')).toHaveCount(0);
|
||||
const places = await page.locator('#feedlist .place b').allTextContents();
|
||||
expect(places.indexOf('Currently Listening')).toBe(places.indexOf('Popular') + 1);
|
||||
|
||||
await page.locator('#feedlist .place', { hasText: 'Currently Listening' }).click();
|
||||
const row = page.locator('#listening .childrow', { hasText: 'Second Episode' });
|
||||
await expect(row).toBeVisible({ timeout: 20_000 });
|
||||
await expect(row).toContainText('14:18 left');
|
||||
|
||||
// Removing it forgets where you got to, so it is still gone on the next visit.
|
||||
await row.locator('[data-a=remove]').click();
|
||||
await expect(row).toHaveCount(0);
|
||||
await expect(page.locator('#player')).not.toBeVisible();
|
||||
await page.locator('#feedlist .place', { hasText: 'Currently Listening' }).click();
|
||||
await expect(page.locator('#listening')).not.toContainText('Second Episode', { timeout: 20_000 });
|
||||
|
||||
// Started again, it is back, and clicking the row resumes it. Finishing it (90%) is
|
||||
// covered in the Rust tests; here the player's own save on close would race it.
|
||||
await page.evaluate(() =>
|
||||
api('/api/entries/test-show/ui-2/position', { method: 'POST', body: JSON.stringify({ secs: 42 }) }));
|
||||
await page.locator('#feedlist .place', { hasText: 'Currently Listening' }).click();
|
||||
await expect(row).toBeVisible({ timeout: 20_000 });
|
||||
await row.click();
|
||||
await expect(page.locator('#player')).toBeVisible();
|
||||
await expect(page.locator('#ptitle')).toHaveText('Second Episode');
|
||||
// The row in the player carries the EQ bars, as the feed view's does, until the player closes.
|
||||
await expect(row).toHaveClass(/\bnow\b/);
|
||||
await expect(row.locator('.eq')).toBeVisible();
|
||||
await page.locator('#pclose').click();
|
||||
await expect(row).not.toHaveClass(/\bnow\b/);
|
||||
await expect(row.locator('.eq')).toBeHidden();
|
||||
});
|
||||
|
||||
test('the filter tabs change what is listed', async ({ page }) => {
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
await expect(page.locator('.ep').first()).toBeVisible({ timeout: 20_000 });
|
||||
const all = await page.locator('.ep').count(); // All is the default tab
|
||||
await expect(page.locator('#count')).toContainText('item');
|
||||
@@ -142,12 +209,12 @@ test('the filter tabs change what is listed', async ({ page }) => {
|
||||
await page.locator('.tabs button', { hasText: 'Unread' }).first().click();
|
||||
expect(await page.locator('.ep').count()).toBeLessThanOrEqual(all);
|
||||
|
||||
await page.locator('.tabs button', { hasText: 'Flagged' }).first().click();
|
||||
await page.locator('.tabs button', { hasText: 'Pinned' }).first().click();
|
||||
await expect(page.locator('#count')).toContainText('0 items');
|
||||
});
|
||||
|
||||
test('a feed URL is editable and has a copy button', async ({ page }) => {
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
await page.locator('#content .acts [data-a="settings"]').click();
|
||||
await expect(page.locator('#surl')).toHaveValue(/show\.xml/);
|
||||
await expect(page.locator('#scopy')).toBeVisible();
|
||||
@@ -187,9 +254,11 @@ test('an OPML subscription is a collapsible folder', async ({ page }) => {
|
||||
body: JSON.stringify({ feed: 'test-subscriptions', force: true }),
|
||||
}));
|
||||
|
||||
// Every row reserves the chevron slot for alignment; only a folder's is clickable.
|
||||
// Only a folder has a triangle, and it is a button that says whether the folder is open.
|
||||
const chev = page.locator('.feed.group .chev');
|
||||
await expect(chev).toBeVisible({ timeout: 20_000 });
|
||||
await expect(page.locator('.feed:not(.group) .chev')).toHaveCount(0);
|
||||
await expect(chev).toHaveAttribute('aria-expanded', 'false');
|
||||
|
||||
// Closed by default: the children are not listed until the folder is opened.
|
||||
const before = await page.locator('.feed').count();
|
||||
@@ -286,7 +355,7 @@ test('opening an item marks it read, and the toggle flips it back', async ({ pag
|
||||
const errors = [];
|
||||
page.on('pageerror', e => errors.push(e.message));
|
||||
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
const row = () => page.locator('.ep', { hasText: 'Second Episode' });
|
||||
await expect(row()).toBeVisible({ timeout: 20_000 });
|
||||
|
||||
@@ -305,7 +374,7 @@ test('opening an item marks it read, and the toggle flips it back', async ({ pag
|
||||
});
|
||||
|
||||
test('the toolbar acts on the selected item', async ({ page }) => {
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
const row = () => page.locator('.ep', { hasText: 'Second Episode' });
|
||||
await expect(row()).toBeVisible({ timeout: 20_000 });
|
||||
// Nothing selected, nothing to act on.
|
||||
@@ -346,6 +415,10 @@ test('a second person has their own feeds and their own read state', async ({ br
|
||||
const ctx = await browser.newContext();
|
||||
const page = await ctx.newPage();
|
||||
await page.goto('/login');
|
||||
// The sign-in page shows the icon, so it has to load before anyone has signed in.
|
||||
const icon = await page.request.get('/icon.png');
|
||||
expect(icon.status()).toBe(200);
|
||||
expect(icon.headers()['content-type']).toBe('image/png');
|
||||
await page.locator('#name').fill('sam');
|
||||
await page.locator('#pw').fill('sampassword');
|
||||
await page.locator('button[type=submit]').click();
|
||||
@@ -353,7 +426,15 @@ test('a second person has their own feeds and their own read state', async ({ br
|
||||
|
||||
// Sam subscribes to nothing yet, so sees nothing -- the admin's feeds are not theirs.
|
||||
await expect(page.locator('#feedlist')).toContainText('No feeds.');
|
||||
await expect(page.locator('#prefs')).toBeHidden(); // not an admin
|
||||
// Settings stays: Sam has their own subscriptions to export and import, and the
|
||||
// schedule and quota are worth seeing even without a say in them. Only the log and the
|
||||
// users screen -- and the server -- are an admin's alone.
|
||||
await expect(page.locator('#prefs')).toBeVisible();
|
||||
await page.locator('#prefs').click();
|
||||
await expect(page.locator('#modalCard')).toContainText('Subscriptions');
|
||||
await expect(page.locator('#gsave')).toBeHidden();
|
||||
await expect(page.locator('#gusers')).toBeHidden();
|
||||
await page.locator('#modalCard .cardacts .btn').first().click();
|
||||
// Hiding the button is not the guard; the server is.
|
||||
expect((await page.request.get('/api/users')).status()).toBe(403);
|
||||
await expect(page.locator('#logs')).toBeHidden();
|
||||
@@ -381,7 +462,7 @@ test('a second person has their own feeds and their own read state', async ({ br
|
||||
|
||||
test('deleting a shared file warns that it is everyone\'s copy', async ({ page }) => {
|
||||
// Admin and Sam both subscribe to Test Show by now, and the daemon downloaded a file.
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
await page.locator('.tabs button', { hasText: 'Downloaded' }).click();
|
||||
const row = page.locator('.ep').first();
|
||||
await expect(row).toBeVisible({ timeout: 20_000 });
|
||||
@@ -405,7 +486,7 @@ test('deleting a shared file warns that it is everyone\'s copy', async ({ page }
|
||||
expect(seen[1]).toContain('one copy of this file');
|
||||
|
||||
await page.reload();
|
||||
await page.getByText('Test Show').click();
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
await page.locator('.tabs button', { hasText: 'Downloaded' }).click();
|
||||
await expect(page.locator('.ep').first()).toBeVisible({ timeout: 20_000 });
|
||||
});
|
||||
@@ -427,6 +508,9 @@ test('an admin adds someone, makes them an admin, and removes them', async ({ pa
|
||||
await page.locator('#uadd').click();
|
||||
const row = userRow(page, 'pat');
|
||||
await expect(row).toBeVisible();
|
||||
// When each account was added and last signed in; the admin signed in with the token link.
|
||||
await expect(row).toContainText(/Added .* never signed in/);
|
||||
await expect(userRow(page, 'admin')).toContainText(/signed in \d+m ago/);
|
||||
await expect(row.locator('[data-a="admin"]')).not.toBeChecked();
|
||||
|
||||
await row.locator('[data-a="admin"]').check();
|
||||
@@ -607,8 +691,8 @@ test('Popular lists what everyone here reads, but never a private feed', async (
|
||||
await piper.locator('#feedlist .place', { hasText: 'Popular' }).click();
|
||||
const offered = piper.locator('#popular .childrow');
|
||||
await expect(offered.filter({ hasText: 'Test Show' })).toBeVisible({ timeout: 20_000 });
|
||||
// An OPML's own feeds ride on the OPML, and a key in a URL marks someone's paid feed.
|
||||
await expect(offered.filter({ hasText: /Grouped Show|grouped-show/ })).toHaveCount(0);
|
||||
// An OPML is listed as the feeds inside it, and a key in a URL marks someone's paid feed.
|
||||
await expect(offered.filter({ hasText: /Test Subscriptions/ })).toHaveCount(0);
|
||||
await expect(offered.filter({ hasText: /Paid Show|paid-show/ })).toHaveCount(0);
|
||||
|
||||
// No URL reaches the page at all, so neither can a key, and the server holds the same line.
|
||||
@@ -617,26 +701,59 @@ test('Popular lists what everyone here reads, but never a private feed', async (
|
||||
expect(listed).not.toContain('.xml');
|
||||
expect((await piper.request.post('/api/popular/paid-show')).status()).toBe(400);
|
||||
|
||||
// Popular is the top ten of the directory, and the directory is every listed feed, A to Z.
|
||||
// Popular is the top ten of the directory, and the directory is every listed feed A to Z,
|
||||
// with an OPML's feeds in place of the OPML in both.
|
||||
const dir = await (await piper.request.get('/api/directory')).json();
|
||||
const top = await (await piper.request.get('/api/popular')).json();
|
||||
const names = dir.map(p => (p.title || p.id).toLowerCase());
|
||||
expect(names).toEqual([...names].sort());
|
||||
const ids = dir.map(p => p.id);
|
||||
expect(ids).not.toContain('test-subscriptions');
|
||||
expect(ids).toEqual(expect.arrayContaining(['grouped-show', 'aardvark-radio']));
|
||||
expect(top.length).toBe(Math.min(10, dir.length));
|
||||
expect(top.every(t => dir.some(d => d.id === t.id))).toBe(true);
|
||||
expect(dir.map(p => p.id)).not.toContain('paid-show');
|
||||
expect(top.every(t => ids.includes(t.id))).toBe(true);
|
||||
expect(ids).not.toContain('paid-show');
|
||||
|
||||
// Subscribe from the directory this time; the popular list shares the same rows.
|
||||
// Subscribe from the directory this time: the same feeds, as a grid of cover art.
|
||||
const tiles = piper.locator('#popular .tile');
|
||||
const pick = (row, name) => piper.locator(`#dirbar .${row} button`, { hasText: new RegExp(`^${name}$`) });
|
||||
await piper.locator('#feedlist .place', { hasText: 'Directory' }).click();
|
||||
await expect(piper.locator('#count')).toContainText(`Directory: ${dir.length} feed`);
|
||||
await expect(offered.filter({ hasText: 'Test Show' })).toBeVisible();
|
||||
await expect(offered.filter({ hasText: /Paid Show|paid-show/ })).toHaveCount(0);
|
||||
await expect(tiles.filter({ hasText: 'Test Show' })).toBeVisible();
|
||||
await expect(tiles.filter({ hasText: /Grouped Show|grouped-show/ })).toBeVisible();
|
||||
await expect(tiles.filter({ hasText: /Test Subscriptions/ })).toHaveCount(0);
|
||||
await expect(tiles.filter({ hasText: /Paid Show|paid-show/ })).toHaveCount(0);
|
||||
await expect(tiles).toHaveCount(dir.length);
|
||||
|
||||
// Two filters that combine: what a feed is, and what it is about.
|
||||
expect(dir.find(p => p.id === 'test-show')).toMatchObject({ podcast: true, category: 'Technology' });
|
||||
expect(dir.find(p => p.id === 'picture-blog')).toMatchObject({ podcast: false });
|
||||
await pick('tabs', 'Blogs').click();
|
||||
await expect(tiles).toHaveCount(dir.filter(p => !p.podcast).length);
|
||||
await expect(tiles.filter({ hasText: 'Test Show' })).toHaveCount(0);
|
||||
// No empty chips: no blog here names Technology, so Blogs does not offer it.
|
||||
await expect(pick('chips', 'Technology')).toHaveCount(0);
|
||||
await pick('tabs', 'Podcasts').click();
|
||||
await pick('chips', 'Technology').click();
|
||||
await expect(pick('chips', 'Technology')).toHaveAttribute('aria-pressed', 'true');
|
||||
await expect(tiles).toHaveCount(dir.filter(p => p.podcast && p.category === 'Technology').length);
|
||||
// A second press lifts the chip and leaves the kind as it was.
|
||||
await pick('chips', 'Technology').click();
|
||||
await expect(tiles).toHaveCount(dir.filter(p => p.podcast).length);
|
||||
await pick('tabs', 'All').click();
|
||||
await expect(tiles).toHaveCount(dir.length);
|
||||
|
||||
// Add a feed opened over Directory fills its own list, not the pane behind it.
|
||||
await piper.locator('#addFeed').click();
|
||||
await expect(piper.locator('#modalCard .childrow', { hasText: 'Test Show' })).toBeVisible();
|
||||
await expect(tiles).toHaveCount(dir.length);
|
||||
await piper.locator('#modalCard button[title="Cancel"]').click();
|
||||
|
||||
const row = async () =>
|
||||
(await (await piper.request.get('/api/popular')).json()).find(p => p.id === 'test-show');
|
||||
const before = await row();
|
||||
expect(before.subscribed).toBe(false);
|
||||
await offered.filter({ hasText: 'Test Show' }).locator('button[title="Subscribe"]').click();
|
||||
await tiles.filter({ hasText: 'Test Show' }).locator('button[title="Subscribe"]').click();
|
||||
await expect(piper.locator('#feedlist .feed', { hasText: 'Test Show' })).toBeVisible({ timeout: 20_000 });
|
||||
|
||||
// Everyone counts, you included: it stays listed, marked as yours, with one more subscriber.
|
||||
@@ -682,13 +799,14 @@ test('a deleted file looks as if it was never downloaded', async ({ page }) => {
|
||||
});
|
||||
|
||||
test('one action, one icon: the toolbar, the page and every dialog agree', async ({ page }) => {
|
||||
const icon = loc => loc.locator('svg path').first().getAttribute('d');
|
||||
// The whole glyph, not just its path: pinned and not pinned share one outline and differ in fill.
|
||||
const icon = loc => loc.locator('svg').first().innerHTML();
|
||||
await page.locator('#feedlist .feed', { hasText: 'Test Show' }).first().click();
|
||||
|
||||
// Unsubscribe is a minus in the toolbar and the feed header, never the x that closes things.
|
||||
expect(await icon(page.locator('#content .acts [data-a="rm"]'))).toBe(await icon(page.locator('#tbRemove')));
|
||||
|
||||
// The toolbar's read and keep show the selected item's state, as its own buttons do, and follow
|
||||
// The toolbar's read and pin show the selected item's state, as its own buttons do, and follow
|
||||
// a change made from the toolbar.
|
||||
await page.locator('.ep').first().click();
|
||||
const pair = async a => [await icon(page.locator(a === 'read' ? '#tbRead' : '#tbFlag')),
|
||||
@@ -768,6 +886,22 @@ test('the item table sorts by any column, both ways, and remembers', async ({ pa
|
||||
await expect(page.locator('#eps .ep .file', { hasText: /\d/ })).toHaveCount(0);
|
||||
});
|
||||
|
||||
test('the selected feed and tab are remembered across a reload', async ({ page }) => {
|
||||
await page.locator('#feedlist .feed', { hasText: 'Test Show' }).first().click();
|
||||
await page.locator('.tabs button', { hasText: 'Unread' }).click();
|
||||
await expect(page.locator('.tabs button.on')).toHaveText('Unread');
|
||||
|
||||
await page.reload();
|
||||
await expect(page.locator('#content h2')).toHaveText('Test Show');
|
||||
await expect(page.locator('.tabs button.on')).toHaveText('Unread');
|
||||
|
||||
// A feed that is gone -- unsubscribed, or never visited on this browser -- lands on All
|
||||
// Subscriptions, not the first feed alphabetically.
|
||||
await page.evaluate(() => localStorage.setItem('ipx.feed', 'no-such-feed'));
|
||||
await page.reload();
|
||||
await expect(page.locator('#feedlist .place.sel')).toContainText('All Subscriptions');
|
||||
});
|
||||
|
||||
test('play in the Files pane plays once, in the player bar', async ({ page }) => {
|
||||
// Regression: the pane had an <audio> of its own, and playing it started the player bar too,
|
||||
// so the same file played twice at once.
|
||||
@@ -775,6 +909,89 @@ test('play in the Files pane plays once, in the player bar', async ({ page }) =>
|
||||
await page.locator('.ep', { has: page.locator('.kind.here') }).first().click();
|
||||
await page.locator('#files [data-a="play"]').click();
|
||||
await expect(page.locator('#player')).toBeVisible();
|
||||
await expect(page.locator('audio')).toHaveCount(1); // the player bar's, and nothing else
|
||||
// The player bar's element doubles as a <video> so a video file has somewhere to show its
|
||||
// picture (see #audio's own comment), but there is still exactly one of it, and nothing else.
|
||||
await expect(page.locator('#audio')).toHaveCount(1);
|
||||
await page.locator('#pclose').click();
|
||||
});
|
||||
|
||||
test('someone the proxy signs in never sees the password page, and signs out through the proxy', async ({ page, browser }) => {
|
||||
// Signed in with the token, not by the proxy: Sign out stays ipx's own.
|
||||
expect((await (await page.request.get('/api/me')).json()).sign_out).toBeNull();
|
||||
|
||||
const ctx = await browser.newContext({ extraHTTPHeaders: { 'X-Test-User': 'proxied@example.com' } });
|
||||
const proxied = await ctx.newPage();
|
||||
// Regression: after Sign out, the password form showed to someone the proxy still vouched for.
|
||||
await proxied.goto('/login');
|
||||
await expect(proxied).toHaveURL(/:8791\/$/);
|
||||
await expect(proxied.locator('#who')).toContainText('proxied@example.com');
|
||||
expect(await (await proxied.request.get('/api/me')).json())
|
||||
.toMatchObject({ name: 'proxied@example.com', sign_out: '/signed-out-by-the-proxy' });
|
||||
await proxied.locator('#signout').click();
|
||||
await expect(proxied).toHaveURL(/\/signed-out-by-the-proxy$/);
|
||||
await ctx.close();
|
||||
});
|
||||
|
||||
test('keys move through items and places, after Feedly', async ({ page }) => {
|
||||
// Last in the file: selecting an item marks it read, which would change what later tests see.
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
const rows = page.locator('#eps .ep');
|
||||
await expect(rows.nth(1)).toBeVisible({ timeout: 20_000 });
|
||||
const sel = page.locator('#eps .ep.sel');
|
||||
const guid = i => rows.nth(i).getAttribute('data-guid');
|
||||
await page.keyboard.press('j');
|
||||
await expect(sel).toHaveAttribute('data-guid', await guid(0));
|
||||
await page.keyboard.press('j');
|
||||
await expect(sel).toHaveAttribute('data-guid', await guid(1));
|
||||
await page.keyboard.press('k');
|
||||
await expect(sel).toHaveAttribute('data-guid', await guid(0));
|
||||
|
||||
// g and a letter go somewhere; typed into a box, the same letters are only text.
|
||||
await page.keyboard.press('g');
|
||||
await page.keyboard.press('d');
|
||||
await expect(page.locator('#count')).toContainText('Directory');
|
||||
await page.locator('#epSearch').focus();
|
||||
await page.keyboard.type('ga');
|
||||
await expect(page.locator('#count')).toContainText('Directory');
|
||||
await page.locator('#epSearch').fill('');
|
||||
await page.locator('#epSearch').blur();
|
||||
|
||||
await page.keyboard.press('?');
|
||||
await expect(page.locator('#modalCard')).toContainText('Keyboard shortcuts');
|
||||
await page.keyboard.press('Escape');
|
||||
await page.keyboard.press('g');
|
||||
await page.keyboard.press('a');
|
||||
await expect(page.locator('#count')).toContainText('All Subscriptions');
|
||||
});
|
||||
|
||||
test('an admin can give a blog its Directory category', async ({ page }) => {
|
||||
const patch = (id, category) => page.evaluate(([id, category]) =>
|
||||
api(`/api/feeds/${id}`, { method: 'PATCH', body: JSON.stringify({ category }) }), [id, category]);
|
||||
const listed = async id => (await page.evaluate(() => api('/api/directory'))).find(p => p.id === id);
|
||||
|
||||
await patch('picture-blog', 'Visual Arts');
|
||||
expect(await listed('picture-blog')).toMatchObject({ podcast: false, category: 'Visual Arts' });
|
||||
// A feed's own iTunes category wins over one given here.
|
||||
await patch('test-show', 'Comedy');
|
||||
expect((await listed('test-show')).category).toBe('Technology');
|
||||
|
||||
for (const id of ['picture-blog', 'test-show']) await patch(id, null);
|
||||
expect((await listed('picture-blog')).category).toBeNull();
|
||||
});
|
||||
|
||||
test('the pinned heading sits over its pins, and the page is set in Inter', async ({ page }) => {
|
||||
await page.locator('.feed', { hasText: 'Test Show' }).click();
|
||||
await expect(page.locator('#eps .ep').first()).toBeVisible({ timeout: 20_000 });
|
||||
const boxes = {
|
||||
headCell: await page.locator('.ephead [data-sort="kept"]').boundingBox(),
|
||||
headIcon: await page.locator('.ephead [data-sort="kept"] svg').first().boundingBox(),
|
||||
rowCell: await page.locator('#eps .ep .fl').first().boundingBox(),
|
||||
rowIcon: await page.locator('#eps .ep .fl svg').first().boundingBox(),
|
||||
};
|
||||
const mid = b => b.x + b.width / 2;
|
||||
expect(Math.abs(mid(boxes.headIcon) - mid(boxes.rowIcon)), JSON.stringify(boxes)).toBeLessThan(1);
|
||||
|
||||
// From ipx itself, not a font service.
|
||||
expect((await page.request.get('/inter.woff2')).headers()['content-type']).toBe('font/woff2');
|
||||
expect(await page.evaluate(() => document.fonts.ready.then(() => document.fonts.check('14px Inter')))).toBe(true);
|
||||
});
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
<rss version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd">
|
||||
<channel><title>Test Show</title><link>http://127.0.0.1:8792/</link><description>A fixture feed.</description>
|
||||
<itunes:image href="http://127.0.0.1:8792/art.png"/>
|
||||
<itunes:category text="Technology"/>
|
||||
<item><title>First Episode</title><guid>ui-1</guid>
|
||||
<pubDate>Mon, 01 Sep 2026 10:00:00 +0000</pubDate>
|
||||
<description><p>Show notes for the first one.</p></description>
|
||||
|
||||
@@ -37,6 +37,10 @@ enabled = false
|
||||
enabled = true
|
||||
bind = "127.0.0.1:8791"
|
||||
token = "${TOKEN}"
|
||||
# The proxy path, for tests that send the header themselves: the daemon sees them at 127.0.0.1.
|
||||
trusted_header = "X-Test-User"
|
||||
trusted_proxies = ["127.0.0.1"]
|
||||
sign_out_url = "/signed-out-by-the-proxy"
|
||||
|
||||
[feeds.test-show]
|
||||
url = "http://127.0.0.1:8792/show.xml"
|
||||
|
||||
92
web/Inter-LICENSE.txt
Normal file
92
web/Inter-LICENSE.txt
Normal file
@@ -0,0 +1,92 @@
|
||||
Copyright (c) 2016 The Inter Project Authors (https://github.com/rsms/inter)
|
||||
|
||||
This Font Software is licensed under the SIL Open Font License, Version 1.1.
|
||||
This license is copied below, and is also available with a FAQ at:
|
||||
http://scripts.sil.org/OFL
|
||||
|
||||
-----------------------------------------------------------
|
||||
SIL OPEN FONT LICENSE Version 1.1 - 26 February 2007
|
||||
-----------------------------------------------------------
|
||||
|
||||
PREAMBLE
|
||||
The goals of the Open Font License (OFL) are to stimulate worldwide
|
||||
development of collaborative font projects, to support the font creation
|
||||
efforts of academic and linguistic communities, and to provide a free and
|
||||
open framework in which fonts may be shared and improved in partnership
|
||||
with others.
|
||||
|
||||
The OFL allows the licensed fonts to be used, studied, modified and
|
||||
redistributed freely as long as they are not sold by themselves. The
|
||||
fonts, including any derivative works, can be bundled, embedded,
|
||||
redistributed and/or sold with any software provided that any reserved
|
||||
names are not used by derivative works. The fonts and derivatives,
|
||||
however, cannot be released under any other type of license. The
|
||||
requirement for fonts to remain under this license does not apply
|
||||
to any document created using the fonts or their derivatives.
|
||||
|
||||
DEFINITIONS
|
||||
"Font Software" refers to the set of files released by the Copyright
|
||||
Holder(s) under this license and clearly marked as such. This may
|
||||
include source files, build scripts and documentation.
|
||||
|
||||
"Reserved Font Name" refers to any names specified as such after the
|
||||
copyright statement(s).
|
||||
|
||||
"Original Version" refers to the collection of Font Software components as
|
||||
distributed by the Copyright Holder(s).
|
||||
|
||||
"Modified Version" refers to any derivative made by adding to, deleting,
|
||||
or substituting -- in part or in whole -- any of the components of the
|
||||
Original Version, by changing formats or by porting the Font Software to a
|
||||
new environment.
|
||||
|
||||
"Author" refers to any designer, engineer, programmer, technical
|
||||
writer or other person who contributed to the Font Software.
|
||||
|
||||
PERMISSION AND CONDITIONS
|
||||
Permission is hereby granted, free of charge, to any person obtaining
|
||||
a copy of the Font Software, to use, study, copy, merge, embed, modify,
|
||||
redistribute, and sell modified and unmodified copies of the Font
|
||||
Software, subject to the following conditions:
|
||||
|
||||
1) Neither the Font Software nor any of its individual components,
|
||||
in Original or Modified Versions, may be sold by itself.
|
||||
|
||||
2) Original or Modified Versions of the Font Software may be bundled,
|
||||
redistributed and/or sold with any software, provided that each copy
|
||||
contains the above copyright notice and this license. These can be
|
||||
included either as stand-alone text files, human-readable headers or
|
||||
in the appropriate machine-readable metadata fields within text or
|
||||
binary files as long as those fields can be easily viewed by the user.
|
||||
|
||||
3) No Modified Version of the Font Software may use the Reserved Font
|
||||
Name(s) unless explicit written permission is granted by the corresponding
|
||||
Copyright Holder. This restriction only applies to the primary font name as
|
||||
presented to the users.
|
||||
|
||||
4) The name(s) of the Copyright Holder(s) or the Author(s) of the Font
|
||||
Software shall not be used to promote, endorse or advertise any
|
||||
Modified Version, except to acknowledge the contribution(s) of the
|
||||
Copyright Holder(s) and the Author(s) or with their explicit written
|
||||
permission.
|
||||
|
||||
5) The Font Software, modified or unmodified, in part or in whole,
|
||||
must be distributed entirely under this license, and must not be
|
||||
distributed under any other license. The requirement for fonts to
|
||||
remain under this license does not apply to any document created
|
||||
using the Font Software.
|
||||
|
||||
TERMINATION
|
||||
This license becomes null and void if any of the above conditions are
|
||||
not met.
|
||||
|
||||
DISCLAIMER
|
||||
THE FONT SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
|
||||
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO ANY WARRANTIES OF
|
||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT
|
||||
OF COPYRIGHT, PATENT, TRADEMARK, OR OTHER RIGHT. IN NO EVENT SHALL THE
|
||||
COPYRIGHT HOLDER BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,
|
||||
INCLUDING ANY GENERAL, SPECIAL, INDIRECT, INCIDENTAL, OR CONSEQUENTIAL
|
||||
DAMAGES, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
|
||||
FROM, OUT OF THE USE OR INABILITY TO USE THE FONT SOFTWARE OR FROM
|
||||
OTHER DEALINGS IN THE FONT SOFTWARE.
|
||||
BIN
web/InterVariable.woff2
Normal file
BIN
web/InterVariable.woff2
Normal file
Binary file not shown.
788
web/index.html
788
web/index.html
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user