Scheduled searches with a weekly review of new and rising repositories

A search can now be saved with a schedule (every N hours, daily, or weekly at
a given day and time). An internal scheduler runs it and keeps a snapshot of
the results on a /data volume (JSON files, no new dependency). The new Watch
tab lists the scheduled searches; the review page of each one shows, for any
run of its history, the repositories never seen before and the ones that
gained the most stars since the previous run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
claude BotandClaude Opus 5.5 committed 2026-10-01 16:36:50 +02:00
1 parent 8f4048d046
commit 02b51fed77
16 files changed
+1745 -45

No files matched your search

+30 -3
View File
@@ -6,7 +6,8 @@ Search GitHub repositories with criteria and browse the results in a clean web i
```
browser ──/api/search──▶ searchgit (Go) ──REST──▶ api.github.com/search/repositories
└── 5 min cache
──/api/saved───▶ ├── 5 min cache
└── scheduler ──▶ /data (JSON snapshots)
```
## Getting started
@@ -18,7 +19,7 @@ docker compose up -d --build
Then open <http://localhost:8080>.
Without Docker: `go run .` (Go 1.23 or newer), then open <http://localhost:8080>.
Without Docker: `go run .` (Go 1.24 or newer, data in `./data`), then open <http://localhost:8080>.
## Criteria
@@ -42,6 +43,24 @@ shows the exact query sent to GitHub (click it to open the same search on github
GitHub returns at most the first 1,000 results of a search.
## Scheduled searches (weekly review)
**Schedule** saves the current filters as a search that the server runs on its own:
every week (day and time), every day, or every N hours, in the server time zone (`TZ`).
Each run keeps a snapshot of the first 30, 50 or 100 results.
The **Watch** tab lists the scheduled searches with their last run, and a review page
per search shows, for any run of its history:
- **New**: repositories never returned by a previous run of this search;
- **Rising**: repositories that gained the most stars since the previous run;
- **All**: the whole snapshot, with the star gain of each repository.
A search runs once as soon as it is created: that first run is the baseline the next
ones are compared with. A run missed while the server was stopped is made at startup.
For a weekly "gems" review, a good start is: a topic or a language, *created < 1 month*
(or *< 1 week*), sorted by *most stars*.
## Configuration
| Variable | Default | Role |
@@ -51,12 +70,20 @@ GitHub returns at most the first 1,000 results of a search.
| `CACHE_TTL` | `5m` | identical searches are served from memory for this long |
| `AUTH_USER` / `AUTH_PASS` | empty | HTTP basic authentication (except `/healthz`) |
| `GITHUB_API` | `https://api.github.com` | API base URL (GitHub Enterprise) |
| `DATA_DIR` | `data` (`/data` in Docker) | scheduled searches and their snapshots (JSON files) |
| `TZ` | `Europe/Paris` in compose | time zone of the schedules |
| `HISTORY_KEEP` | `100` | runs kept per scheduled search (0 = all) |
## API
- `GET /api/search?q=&in=&language=&stars=&maxstars=&minforks=&topic=&license=&pushed=&created=&goodfirst=1&forks=1&archived=1&sort=&order=&page=&per_page=`
(`pushed` / `created`: `1w`, `1m`, `3m`, `6m`, `1y`, `2y`, `5y`)
- `GET /api/status`: token configured and last known search quota
- `GET /api/status`: token configured, last known search quota, server time zone
- `GET|POST /api/saved`, `GET|PUT|DELETE /api/saved/{id}`: scheduled searches
(`{"name", "params", "maxResults", "enabled", "schedule": {"every": "week|day|hours", "weekday", "hour", "minute", "hours"}}`,
`params` being the `/api/search` query string)
- `POST /api/saved/{id}/run`: run now
- `GET /api/saved/{id}/runs`, `GET /api/saved/{id}/runs/{run|latest}`: history and snapshots
- `GET /healthz`
## Development