- Not Reddit (a forum, a reviews API, an internal dataset) → write a source. You map your items to the standard shape and they work like any built-in source. See Build a source.
- Reddit data from somewhere other than the public archive (local files, your
own database, a cache) → implement
CorpusReaderandCommentSource, two small interfaces that hand metalworks raw post and comment rows. That’s this page.
reddit source normally reads from a public Reddit archive. Implement these
two interfaces and it reads from your data instead.
The protocols
RedditPost / RedditComment
and writes them to your store.
Wire it in
Comments are optional
If you have submissions but no comment source, passcomments=None (the default
when offline). The pipeline marks the report partial with a caveat rather than
failing. Cluster quotes come from comments, so a comments-less run produces
submission-level signal only.
Fully offline
Point a local reader at committed parquet and use fake models and an in-memory store, and the whole pipeline runs with no network. This is the pattern for tests and air-gapped runs.Data that isn’t Reddit
For anything that isn’t Reddit (a forum, reviews, your own dataset), don’t useCorpusReader — build a source instead. You map your items
to the standard shape and they work like any built-in source. There are three lanes
to pick from: a grounding connector that yields quotable records, a magnitude
provider that attaches a number (downloads, search volume) to a theme, or an agentic
discovery provider that reaches the long tail.