Retention panel's "currently using" line and usage bar were always blank
— totalStorageBytes() existed but was never wired into the settings
response, so retention.storageUsedMB was undefined on every load. Now
computed fresh on every GET/PATCH /api/admin/settings.
Tracked events gain a keywords field: only items whose title/summary/
body contain at least one of them (word, phrase, or emoji — e.g. 🇮🇷 for
an "Iran war" event) qualify for that event's recap, instead of every
item from its assigned sources. An item from an event-linked source that
doesn't match now falls through to normal synthesis rather than being
silently dropped. Also added the source-assignment + keyword-filter edit
UI to EventsTab.svelte, which had no way to populate sourceIds at all
before this (the "assign from the Sources tab" comment referenced a
feature that was never built).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Was backwards — the card displayed the polled channel's own name/avatar
with a "Forwarded from @origin" line. Now mirrors the tweet retweet
pattern exactly: the card's author identity (name/handle/avatar) is
always the original channel/user, forward or not, same as tweet.authorName
never being the retweeter. A new repostedByHandle field (replacing
forwardedFrom) carries the polled channel's own handle for the
"Forwarded by @X" line above the card.
Media resolution still keys off the polled channel specifically (a new
sourceChannelUsername field) since attached media lives on the polled
channel's own copy of the message regardless of who originally posted it
— only the avatar now resolves against the displayed (possibly origin,
possibly null) channel identity.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Mirrors the repost-line treatment tweets already get. GramJS resolves a
forward's origin channel/user from entities Telegram sends alongside the
same getMessages response (message.forward.chat/.sender), falling back
to fwdFrom.fromName for the rarer case where the origin hid its identity.
TelegramCard shows "↪️ Forwarded from @username" (or just the name if no
public handle) above the meta row.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Telegram has no public hotlinkable media URL the way Twitter's CDN does,
so there's no "direct" option: self-host downloads via the logged-in
account and stores locally; proxy re-fetches live through that same
account on each view via a new /media/telegram-proxy route (small
in-memory cache to absorb repeat views), without persisting anything.
Adapter no longer downloads media eagerly at ingestion — it only records
lightweight refs (message id, kind, mime type, dimensions); publish.ts
resolves those into a servable url per the admin's chosen mode, same
timing as Nitter's tweet media resolution. New "Telegram (message media)"
panel added to the admin Retention tab alongside the existing Nitter one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Logs into a real Telegram account via MTProto (GramJS) rather than a bot,
so it can read any public channel's history. API ID/hash and the login
session are entered through the admin Connections panel and stored
encrypted at rest (new storage/crypto.ts AES-256-GCM helper) rather than
via .env. Messages render as their own TelegramCard (same treatment as
tweets) and open the original message on Telegram instead of an internal
article page; attached media/albums are downloaded and self-hosted at
ingestion time since Telegram has no public hotlinkable media URL.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Unlike "Clear content" (which deletes both the raw ingested items and
the published articles), reissue keeps the raw content_items and only
deletes the articles built from them, resetting cluster_id so those
items get picked up and republished by the very next scheduler tick.
This is for picking up pipeline/rendering changes on already-ingested
content without depending on the source feed to serve the same items
again — Nitter/Twitter in particular won't reliably resurface an old
tweet on a fresh poll. Same multi-source protection as clearing: an
article merged from this source's items together with another
source's is left alone entirely, since undoing just one contributor's
share of a merge isn't supported.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
extractQuotedTweet() was capping the quoted tweet's text at 240
characters via toSummary(); it now keeps the full cleaned text,
matching the outer tweet's own body which was never truncated.
Dropped the matching -webkit-line-clamp: 3 on .quoted-text in
TweetCard.svelte so the full text actually renders instead of being
clipped after 3 lines. Media sizing (.quoted-img's 140px cap, the
media grid's fixed cell heights) is unchanged — only text was meant
to stay unconstrained.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Nitter's RSS marks a bare retweet with a "RT by @handle:" prefix on
the item's <title> (dc:creator is already the original author, not
the retweeter — confirmed against a real sample). TweetCard now shows
a "🔁 Reposted by @handle" line above an otherwise-unchanged card.
A quote-tweet's RSS description carries the embedded tweet fully
inline in a <blockquote> (author, text, one image, permalink) — no
extra fxtwitter call needed. TweetCard renders it as a smaller frame
nested inside the same outer card, below the quoting tweet's own
text and media, labeled "↩️ Replying to @handle" per how this reads
to a visitor even though it's technically Nitter's quote-tweet
representation. Clicking it opens that tweet's own permalink,
independent of the outer card's link.
Both parsers verified against the real sample RSS (Polymarket/
rawsalerts retweet, Goldman/zerohedge quote-tweet) and against the
live publishDirect pipeline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
authorName came from fxtwitter's enrichment (keyed by the tweet ID, which
for a retweet/reply resolves to the original tweet) while authorHandle was
hardcoded to the RSS feed's dc:creator — which is actually whichever list
member's retweet or reply surfaced the item, not the original author. The
card ended up showing one person's display name next to another person's
@handle. Both fields now come from the same enrichment response, falling
back together to the RSS-derived handle only when enrichment fails.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Categories can now be marked "Private" in the admin panel's Category
priority list. Private categories (and every article tagged with
one, even if it's also tagged with a public category) are hidden
from /api/categories, /api/feed, and /api/article/:id for anyone
without a valid login — a plain visitor's browser, not the admin API
key, since that's a header-based credential for the admin SPA only.
Login is a single shared password set via PRIVATE_ACCESS_PASSWORD in
the backend's .env (unset by default, which disables the feature
entirely). On success the backend sets a stateless httpOnly cookie —
its value is a deterministic hash of the password, checked with a
timing-safe comparison on every request, so there's no session table
to maintain. The cookie is requested at the ~400-day cap browsers
enforce on persistent cookies, the closest a cookie can get to
"retained indefinitely."
On the frontend, an always-visible lock icon in the masthead (shown
whenever the feature is configured, independent of the admin panel's
own enabled/disabled toggle) opens a password prompt and reflects
locked/unlocked state.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Categories can now be marked "Private" in the admin panel's Category
priority list. Private categories (and every article tagged with
one, even if it's also tagged with a public category) are hidden
from /api/categories, /api/feed, and /api/article/:id for anyone
without a valid login — a plain visitor's browser, not the admin API
key, since that's a header-based credential for the admin SPA only.
Login is a single shared password set via PRIVATE_ACCESS_PASSWORD in
the backend's .env (unset by default, which disables the feature
entirely). On success the backend sets a stateless httpOnly cookie —
its value is a deterministic hash of the password, checked with a
timing-safe comparison on every request, so there's no session table
to maintain. The cookie is requested at the ~400-day cap browsers
enforce on persistent cookies, the closest a cookie can get to
"retained indefinitely."
On the frontend, an always-visible lock icon in the masthead (shown
whenever the feature is configured, independent of the admin panel's
own enabled/disabled toggle) opens a password prompt and reflects
locked/unlocked state.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
The whole card previously linked to our own /article/[id] page. Per
request, clicking the card frame now opens the original tweet (in a
new tab) instead — there's no separate "full article" view for a
tweet anyway. Clicking a photo opens that image by itself in a new
tab (still resolved through the configured media mode, so proxy mode
doesn't leak the browser's IP to Twitter when viewing the full image
either). Clicking a video's native controls still just plays/pauses
it rather than navigating anywhere.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Firefox needs an explicit type on <source> to pick a decoder,
especially when the URL's extension is followed by a query string
(?tag=29) rather than ending cleanly in .mp4. fxtwitter's own
response confirms these are always video/mp4.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
article.tweet?.media.slice(...) only guarded against a missing tweet
object, not a missing media array — pre-existing published tweets in
the DB predate that field and threw "Cannot read properties of
undefined (reading 'slice')" on every category page containing one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Tweets can carry up to 4 photos/videos/gifs; fxtwitter's media.all
preserves their original order and, for videos, gives a real playable
.mp4 plus a poster thumbnail. TweetCard now renders these as a
1/2/3/4-item grid (Twitter's own layout shapes) with fixed cell
heights so a tall portrait image no longer dictates the whole card's
height in the column view, and video/gif items play back with native
controls instead of showing a static frame. Each item's url (and a
video's thumbnail) still resolves through the configured Nitter media
mode (self-host/proxy/direct) individually.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
Adds nitterMediaMode (self-host/proxy/direct, default proxy) and
fxtwitterBaseUrl to global settings with a new Retention tab panel.
Tweet images and avatars now resolve through the chosen mode instead
of always being downloaded — proxy mode streams media through a new
SSRF-hardened /media/proxy route (hostname allowlist + DNS-rebinding
defense) so the origin server's IP is never exposed to Twitter's CDN,
direct hotlinks the original URL, and self-host keeps the prior
always-download behavior. fxtwitterBaseUrl lets the enrichment call
target a self-hosted FixTweet mirror instead of the public instance.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014c1L8ghNBFjfiH64UMViP8
/category/x-news filtered by the raw URL slug ("x-news") instead of the
real category name ("X News"), so it never matched merged_articles.category
values for any multi-word category — only worked for the seeded defaults
because they're all single words where the slug and name happen to be
identical once lowercased. Now resolves the slug back to the actual
category name via the site's own category list before filtering.
resolveHeroImage's favicon fallback exists so regular articles never look
entirely bare, but for a tweet it meant an image-less tweet showed the
Nitter instance's own favicon slapped on as if it were the tweet's photo.
publishDirect now skips that fallback specifically for tweet items —
TweetCard.svelte already renders cleanly with no image at all.
Was calling /2/status/<id> (no username) based on the originally-given
example; the actual working endpoint is /<handle>/status/<id> (no version
prefix), confirmed via a real curl response. text, created_timestamp, and
author.name/avatar_url all match the assumed shape exactly — only
media.photos[].url remains unverified (that test tweet had no photo), still
guarded by the existing RSS-image fallback either way.
Nitter list/user RSS feeds are ingested as their own source type, enriched
via fxtwitter (author name/handle/avatar, cleaner text, attached photo) with
a graceful RSS-only fallback when that enrichment fails. Tweets always
publish directly, one per article, and never enter the LLM
clustering/synthesis pipeline — the same bypass already used for YouTube,
since merging unrelated tweets together makes no sense.
Rendering: a new distinct embed-card component (avatar, name + @handle,
full untruncated text, optional attached image, published-date-only
timestamp, no like/retweet stats) replaces the plain article row wherever a
tweet appears, on both the category-page list and the article detail page.
Verified end-to-end against the real sample Nitter RSS feed (served
locally): ingestion (all 100 items, tweet metadata correctly extracted,
retweet/quote-tweet blockquotes correctly excluded from own-content text),
publishing (bypasses clustering, tweet field threaded through to the
published article), and rendering (embed card appears on the homepage feed
and the article detail page, no duplicate title).
Known follow-up: fxtwitter's JSON field names are based on public
documentation, not a verified live response (that API is unreachable from
this sandbox) — worth a real curl check before relying on the enrichment
path in production; the RSS-only fallback path is what's actually been
exercised here.
/category/local secretly filtered by geo:'philadelphia' instead of the
Local category, a leftover from an old "Local: <region>" colon-syntax
convention that has no admin UI behind it anymore — the Sources tab assigns
plain category names via checkboxes, so a source tagged "Local" never got a
matching geo value and could never show up here, even though it correctly
appeared on Top Stories (which only checks pushToTopStories + the category
array, not geo). Local now filters by category like every other category
page.
Verified: reproduced the exact bug (geo: null despite category: ['Local']),
confirmed /api/feed?geo=philadelphia returns nothing while the new
/api/feed?category=local correctly returns the article.
A long RSS/Google News URL with no natural wrap points would overflow its
grid column and break the whole row layout. The URL is still visible (and
editable) via the edit form — the list row now only shows a status line
when there's something to say (poll result, error, or a just-cleared note).
Confirmed by reproduction: a mismatch between backend FRONTEND_ORIGIN and
frontend ORIGIN throws "CORS error: Incorrect 'Access-Control-Allow-Origin'
header is present on the requested resource" on every page, since
SvelteKit's server-side fetch enforces real CORS during SSR. Easy to trip on
since the two values live in separate .env files edited at different times.
adapter-auto doesn't produce a runnable standalone server when it can't
detect a supported hosting platform (Vercel/Netlify/Cloudflare/etc.) — this
is a self-hosted app with no such platform, so builds were silently missing
a real server output. Swapped to adapter-node, which builds to
build/index.js: a persistent Node server that reads PORT/HOST/ORIGIN at
runtime, exactly what a reverse proxy needs to point a domain at.
Added a README section covering the full path: building/running both apps
as plain Node processes, the env vars each needs, and the Nginx Proxy
Manager side (proxy host + Custom Locations for /api and /media under a
single-domain, path-routed setup, or a simpler two-domain alternative).
Verified live: the adapter-node build actually serves pages and correctly
picks up ADMIN_PANEL_ENABLED via `node --env-file=.env build/index.js`
(SvelteKit's $env/dynamic/private reads process.env directly in production,
unlike the vite-dev-time gap from the previous fix).
Only /admin/login and /admin/settings existed as actual pages, so navigating
to /admin itself 404'd regardless of ADMIN_PANEL_ENABLED. Settings already
redirects to login on a 401, so that's the sensible default landing spot.
process.env.ADMIN_PANEL_ENABLED was always undefined in the running
SvelteKit server process — Vite only injects VITE_-prefixed vars into
process.env for server-side code; plain vars in frontend/.env never reached
it, so the admin panel stayed disabled (cog hidden, /admin/* 404s) no matter
what the .env file said. Switched to SvelteKit's own $env/dynamic/private,
which reads it correctly in dev, preview, and adapter-based deployments.
Reproduced and verified the fix against a real frontend/.env file (not an
inline shell var, which is what masked this the first time).
Two hardening changes beyond just a password:
- The admin panel no longer uses stored credentials at all. The backend
generates a random API key on every startup and prints it to its own
console (never through the DB-backed logger, since that's only reachable
from inside the panel this key protects). Every /api/admin/* request must
carry it as an X-Api-Key header, checked with a timing-safe comparison on
every call — there's no session to create or steal, and restarting the
backend invalidates the previous key immediately. The old admin_users and
sessions tables, scrypt password hashing, and cookie-based session plumbing
are removed entirely (dropped via migration for existing installs, not
left behind unused). The login page keeps its existing layout but now asks
for this key and explains where to find it, storing it in the browser's
localStorage rather than relying on a server session.
- The admin panel (the masthead's cog icon and the /admin/* pages
themselves) is now disabled by default on every deployment, gated by a new
frontend-only ADMIN_PANEL_ENABLED env var. This is a separate, UI-only
visibility control — the API key above is what actually protects the
backend regardless of this flag.
Every ingested article used to show up on the homepage regardless of its
source, which meant a handful of high-volume feeds could flood "Top Stories."
Sources now default to not appearing there; a source has to explicitly opt
in via a new checkbox (also toggleable inline with a star icon) for its
articles to show up on the homepage feed. An article shows there if any of
its contributing sources opted in — merged/clustered stories aren't held to
requiring all sources to agree. Category pages, Local, tags, and events are
unaffected; this only gates the bare, no-filter homepage query.
Schema: sources.push_to_top_stories and merged_articles.top_stories, both
backfilled for existing databases via ALTER TABLE.
YouTube's public Atom feed only accepts a channel_id (or the legacy user
param) — it has no equivalent for the newer @handle format, so pasting a
handle URL straight into the source's url field wouldn't have worked. The
adapter now accepts a bare channel ID, a /channel/UC... URL, an @handle URL,
or a bare handle/username, resolving whichever was given to the actual
channel ID by reading it off the channel page when needed.
- Fix "Body cannot be empty" error on DELETE by making the JSON content-type
parser tolerate empty bodies, and by only sending Content-Type from the
frontend when a request actually has one.
- Deleting a source now cascades: raw content items and any article composed
entirely from that source are removed too, plus their media.
- Add admin endpoints/UI to clear all articles, all media, or a single
source's content without deleting the source, so things can be repopulated
fresh.
- Sources can now be assigned multiple categories via checkboxes (instead of
free text) and edited in place, not just added/deleted.
- Add a "News" default category (seeded fresh, backfilled on existing DBs) so
general news sources have a real home instead of the pseudo-category "Top
stories", which is just the homepage's all-categories chronological view.
- Widen the site's content column 15% (1080px -> 1242px).
- Add YouTube as its own source type/ingestion module: pulls a channel's
public Atom feed, and each video always publishes directly as its own
article (title, embedded video, publish date, description) rather than
going through the cross-source clustering/synthesis pipeline.