Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

lightcurves

Measure how a social post's reach rises and decays over time, and classify posts by the shape of that curve.

In astronomy, a light curve is a plot of an object's brightness over time. It's the primary instrument for identifying transients — supernovae, variable stars, exoplanet transits are classified by curve shape, not by peak brightness. A Type Ia supernova is recognisable because its curve rises and falls a particular way.

A social post is a transient. Impressions are luminosity. Every analytics dashboard shows you the peak; almost none of them show you the shape.

This tool collects the shape.

What you are signing up for

Be clear about this before reading further — it is a real commitment, not a script.

A long-running daemon. Not a cron job, and not serverless. Each post is sampled at offsets from its own publish time: a post published at 12:15 is read at 12:45 and 13:15, one published at 15:27 at 15:57 and 16:27. Nothing batches onto a shared clock, because at the front of a decay curve that is where all the shape is. A scheduler that snaps work to a grid destroys the measurement it is taking.

A Postgres database you keep. The data is a time series that cannot be backfilled. If the daemon is off for a day, that day is gone — not delayed. Some sources make this worse: LinkedIn's member changelog retains 28 days, so events older than that are unrecoverable by anyone, including LinkedIn.

API credentials for each source, and the patience to get them. Buffer needs an API key. LinkedIn needs an OAuth app with r_member_postAnalytics, which may require approval.

If you want a number for how a post did, your platform's dashboard already has it and you should use that. This exists for the shape over time, which no dashboard shows.

Requirements

Go 1.25+ (to build; the deployed artifact is a single static binary)
PostgreSQL 14+ — uses generate_series, lateral joins, JSONB
Runtime anything that runs a Linux binary continuously: a VPS, systemd on a box you already have, a container. ~30 MB RSS
Network outbound HTTPS to each source's API
Credentials per source; see docs/SETUP.md

No cloud is assumed and none is required. The tool needs a database and a process, and baking one provider's orchestration into the core would make it worse everywhere else — so the primitives stay primitive: deploy/ has a systemd unit and a Dockerfile, and deploy/setup/ has a provider-neutral, three-step installer that stands the whole thing up against any Postgres. If you happen to be on AWS, deploy/aws/ adds an optional CloudFormation template for a small RDS instance — one you run yourself and can read first. The setup scripts never create cloud resources; at most they discover an RDS instance you already have and ask before using it. That is the line: convenience for a provider you already chose, never orchestration baked into the tool.

The three steps, kept separate so the architecture stays visible — deploy/README.md has the detail:

deploy/setup/01-database.sh     # a database + a no-DELETE role, on any Postgres
deploy/setup/02-application.sh  # the daemon, as a container by default
deploy/setup/03-verify.sh       # confirm the two can actually talk

Why the shape

A post that reaches 2,000 people in six hours and one that reaches 2,000 over a week are different posts with different lessons, and they are indistinguishable in every report that gives you a single number.

The questions that need the curve:

  • Carousels usually reach ~150 in the first three hours. When do they break out instead?
  • If a post isn't sustaining 15 impressions/hour by hour seven, is it finished?
  • Something spiked at 72h. Was there a comment at 72h?

None of these are answerable from a total. All of them are answerable from a time series plus the post's own attributes.

What it stores

Samples — periodic readings of each post's metrics, at intervals you configure, anchored to that post's publish time. A post published at 12:15 gets its readings at 12:45, 13:15, and so on. Nothing is batched onto a shared clock, because at the front of a decay curve that's where all the shape is.

Events — individual comments, reactions and reshares with their own timestamps, where the platform exposes them. Events carry the platform's timestamp rather than the observation time, so they're exact regardless of how often you poll. This is what turns "there was a spike at 72h" into "there was a comment at 72h14m."

Attributes — arbitrary key/value pairs describing the post: format, hook style, whether it had a link, which template produced it. Bring your own vocabulary; the tool doesn't need to know what your dimensions mean, only that you want curves grouped by them.

The post itself — text and metadata, so the dataset is self-contained and analysis never requires joining back to wherever the post was authored.

Grouping posts: campaign_id

posts.campaign_id is a free-text label you assign, so you can ask questions about a set of posts rather than one at a time:

SELECT campaign_id, avg(value)
FROM posts p
JOIN collections c ON c.post_id = p.id
JOIN readings   r ON r.collection_id = c.id
WHERE r.metric_key = 'impressions' AND c.elapsed_s BETWEEN 10800 AND 14400
GROUP BY campaign_id;

Despite the name it is not only for campaigns. It is whatever grouping you care about — a client, a product launch, a content series, a month. If you need more than one kind of grouping at once, use post_attributes instead; campaign_id is the one that gets a column because it is the grouping almost everyone wants and nobody wants to join for.

There is no campaigns table, and that is deliberate. lightcurves does not run campaigns and has no opinion about what one is. Whatever system you already plan in — a spreadsheet, a Notion database, a directory name — owns their identity, and this column just carries the identifier far enough to group by. A foreign key would mean registering a campaign here before a post could be labelled, which turns an optional label into a setup step.

Nothing has to be pushed anywhere. You do not need Buffer tags, or any equivalent feature on any platform. Some tools offer one and they are fine to use, but a platform-side tag only describes posts published through that platform — post something directly and it cannot describe that at all, leaving you attribution that is silently incomplete rather than absent.

The one thing that has to be true is that you can match your own content to the post id a source knows it by — buffer_post_id or platform_urn. Those are unique per row: the same text sent to three networks is three posts with three ids.

Label whenever you like. Setting it at publish time is tidy, and it is what happens if your publishing already knows the grouping. But it is only an UPDATE, so labelling a year of history later works exactly as well:

UPDATE posts SET campaign_id = 'obvious-math-mistake'
WHERE buffer_post_id IN ('6a755487ab236f0ec5886863', '6a7554b58b39b04b79f1f3ff');

The metrics were collected regardless. Grouping is a question you ask afterwards, and nothing about the collection depends on having decided it first.

Longer treatment, with a worked example of a directory-per-campaign layout and what an importer should report before it writes anything: docs/GROUPING.md. The campaign-import skill in .claude/skills/ builds the importer for whatever layout you already have.

Why attributes matter more than volume

Insight lives in comparisons, not in individual posts. n posts give you n(n−1)/2 pairwise comparisons, and k attributes give you up to 2^k ways to slice them.

So the dataset compounds: quadratically in posts, exponentially in attributes. Two hundred posts with twelve well-recorded dimensions will out-answer two thousand posts with none. Record generously and cheaply — you cannot know in advance which slice holds the finding, and you cannot backfill a time series you didn't collect.

Status

Early. Nothing is released yet.

The daemon works: it derives what is outstanding, samples every enabled API version of a slot in one observation, records failures as explainable rows, and picks up discovered posts. Buffer is verified against a live key (2026-08-10): batching, rate-limit header parsing, metric names and metricsUpdatedAt all confirmed against the real API. LinkedIn is not — its access request is pending, so treat that response mapping as unverified until a token exists.

Collectors:

source samples events discovery notes
Buffer yes no yes counts only; refresh appears ~daily. A per-day request quota is the real constraint, and is learned from the response headers rather than configured
LinkedIn yes not yet no 1 request per metric (up to 11), and Development tier allows only 100 calls/member/day. Needs a share/ugcPost URN, which Buffer's permalinks do supply

LinkedIn events — the comment and reaction timestamps that turn "a spike at 72h" into "a comment at 72h14m" — need an endpoint that r_member_postAnalytics does not reach. The collector deliberately does not implement the event interface rather than claiming support and returning nothing.

How it works

docs/ARCHITECTURE.md has the diagrams: how the schedule is derived rather than stored, the two sampling paths and when each is right, the data model, and how the daemon measures things the vendors do not document.

Trying it

Credentials are the part that goes wrong, so there is a command that does nothing but tell you whether yours work:

cp .env.example .env          # fill in your keys; .env is gitignored
go build -o lightcurves ./cmd/lightcurves
./lightcurves -check

It connects to the database, verifies every configured credential with the cheapest authenticated call each source offers, reports rotation deadlines, and exits. The same check runs at startup — where a rejected credential stops the daemon, but an unreachable source does not, because refusing to start during an outage loses curve that cannot be recovered.

Or with containers, which also stands up Postgres with the schema applied:

docker compose -f deploy/compose.yaml run --rm lightcurves -check
docker compose -f deploy/compose.yaml up -d

Scope

Sub-minute sampling is out of scope. One minute is the finest granularity this tool targets. That's a deliberate boundary — it rules out the infrastructure that second-resolution polling would require, and no social platform's numbers move meaningfully faster than that.

License

MIT

About

Samples a social post's reach over time and classifies posts by the shape of the decay curve

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages