Tampon disque, positions Docker sans perte et DNS inverse non bloquant #15

Merged
claude Bot merged 1 commits from feat/fiabilite-ingestion into main 2026-10-03 16:39:40 +02:00

Requested by Cédric

Before: when VictoriaLogs was unreachable, the store loop slept up to 15 s between retries, then dropped the batch. Docker positions moved as soon as a line was read, so a dropped line was never read again. A new sender IP made the UDP listener wait up to 300 ms for its DNS name, and a flood of unknown IPs emptied the whole name cache.

After: batches VictoriaLogs refuses or cannot take are written to /data/spool (up to SPOOL_MAX_MB, 1 GiB by default) and sent again, oldest first, once it answers. The bottom bar shows how many messages wait on disk. Docker and host logs slow down on a full queue instead of losing lines, and the Docker position only moves once a line is stored or spooled. Syslog listeners never wait for DNS.

How:

  • spool.go: one NDJSON file per batch (atomic write), counted on startup, oldest first; a batch VictoriaLogs rejects with a 4xx (other than 429) is dropped rather than retried forever.
  • store.go: while older batches wait on disk, new ones queue behind them (order kept). A replay loop with backoff (1 s to 30 s) drains the spool. Shutdown gives the last batches 5 s, then spools them. SPOOL_MAX_MB=0 keeps the old behaviour (retries, now context-aware).
  • Entry.Wait / Entry.Done: producers that can wait (Docker, host logs) block on a full queue. Done runs once the line is stored or spooled; Docker moves its checkpoint there.
  • rdns.go: LRU cache (10 000 entries) instead of clearing it, at most 64 lookups at once, and Resolve, which calls back from a goroutine (at most 1 024 waiting) instead of blocking the listener.
  • Tests: spool, store (VictoriaLogs down then back, order, refused batch, no spool), queue backpressure, rDNS limits and LRU. go test -race passes.

Note: after a Docker reconnect, lines still in the queue can be read twice (a duplicate rather than a loss).

🤖 Generated with Claude Code

<!-- ccr-projects-attribution: {"github_login":"vblogio"} --> _Requested by **Cédric**_ Before: when VictoriaLogs was unreachable, the store loop slept up to 15 s between retries, then dropped the batch. Docker positions moved as soon as a line was read, so a dropped line was never read again. A new sender IP made the UDP listener wait up to 300 ms for its DNS name, and a flood of unknown IPs emptied the whole name cache. After: batches VictoriaLogs refuses or cannot take are written to `/data/spool` (up to `SPOOL_MAX_MB`, 1 GiB by default) and sent again, oldest first, once it answers. The bottom bar shows how many messages wait on disk. Docker and host logs slow down on a full queue instead of losing lines, and the Docker position only moves once a line is stored or spooled. Syslog listeners never wait for DNS. How: - `spool.go`: one NDJSON file per batch (atomic write), counted on startup, oldest first; a batch VictoriaLogs rejects with a 4xx (other than 429) is dropped rather than retried forever. - `store.go`: while older batches wait on disk, new ones queue behind them (order kept). A replay loop with backoff (1 s to 30 s) drains the spool. Shutdown gives the last batches 5 s, then spools them. `SPOOL_MAX_MB=0` keeps the old behaviour (retries, now context-aware). - `Entry.Wait` / `Entry.Done`: producers that can wait (Docker, host logs) block on a full queue. `Done` runs once the line is stored or spooled; Docker moves its checkpoint there. - `rdns.go`: LRU cache (10 000 entries) instead of clearing it, at most 64 lookups at once, and `Resolve`, which calls back from a goroutine (at most 1 024 waiting) instead of blocking the listener. - Tests: spool, store (VictoriaLogs down then back, order, refused batch, no spool), queue backpressure, rDNS limits and LRU. `go test -race` passes. Note: after a Docker reconnect, lines still in the queue can be read twice (a duplicate rather than a loss). 🤖 Generated with [Claude Code](https://claude.com/claude-code)
claude Bot added 1 commit 2026-10-03 16:26:36 +02:00
- Batches that fail go to /data/spool (SPOOL_MAX_MB, 1 GiB by default) and
  are sent again oldest first; retries no longer block the store loop and
  follow the shutdown context.
- Docker and host logs wait for room in a full queue instead of being
  dropped; the Docker position only moves once a line is stored or spooled.
- Reverse DNS no longer holds up the syslog listeners, with an LRU cache
  and a cap on concurrent lookups.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
claude Bot marked the pull request as ready for review 2026-10-03 16:39:31 +02:00
claude Bot merged commit 524f5dc0fd into main 2026-10-03 16:39:40 +02:00
claude Bot deleted branch feat/fiabilite-ingestion 2026-10-03 16:39:40 +02:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: vLab-BZH/logstream#15