Engineering & Technical Deep-Dives
The bulk of the week's attention went to a 10-post thread walking through six real evaluation patterns PostHog uses to monitor AI quality across millions of LLM generations per week [1]. The thread also included a substantive reply clarifying that review agents can auto-merge fixes depending on complexity and importance, or else write the fix and open a PR for human review [1].
Code-based evals: SQL anti-patterns and sentiment checks: The first eval example checks SQL queries written by PostHog's execute-SQL tool — no LLM involved, just code scanning for known anti-patterns like positive leading wildcards or missing timestamps, so the team can tune prompts to prevent them [1]. A second basic eval checks sentiment of PostHog AI conversations to surface negative traces for investigation [1].
Eval-led development and LLM-as-judge: @_raduraicea used eval-led test-driven development: create an eval that checks previous outputs to avoid repetition, watch it start at ~0% pass rate, ask PostHog in Slack to fix what's failing, and reach 100% pass rate [1]. A separate LLM-as-judge eval checks outreach briefs for unverified names — LLMs use the first name of an email as a person's name about ~7% of the time, and an LLM check catches the errors [1].
Non-binary evals and the self-driving feedback loop: Evals don't have to be pass/fail — the product grounding eval returns N/A about ~90% of the time because most interactions don't make a factual claim at all [1]. Over the last 30 days, evals generated 127 reports and 14 fixes to issues like PostHog AI not creating insights it said it did, the SQL tool failing, and sales briefs inventing people's names [1]. Every eval run is an AI Observability event, with 100k free per month [1]. Separately, in reply to a question, @posthog noted that review agents can auto-merge fixes depending on complexity and importance, or write the fix and open a PR for the team to review [1].
AI & Autonomy Thesis
Continuing the week's AI thread, @posthog linked to a blog post by Andy on the 6-7 loops PostHog uses every day to build self-driving — from agent feedback to scouts that fix their own instructions, with real reports, merged PRs, and notes on where humans still fit [2]. The post notably skipped the usual X article format so readers had to click through, because Andy embedded a playable loops game at the end [2]. A separate article tackled what happens to engineers when AI writes all the code [3]. Over the last 4 months at PostHog, agent-opened PRs in the monorepo went from ~20% to 70%, which breaks the definition of engineer as "someone who writes code" [3]. The argument: engineers won't go extinct — Anthropic itself claims "coding is largely solved" yet has 208 open roles with "engineer" in the title [3]. Instead, engineers shift to "monitoring the situation" — watching agents, fixing errors, responding to customers, checking dashboards, strategizing — and the best ones build systems (loop engineering, context engineering) to get the information they need at the right time [3]. PostHog engineers already have standup bots, repo summary scouts, and custom setups for monitoring in-progress PRs [3].
Developer Marketing & GTM Strategy
@posthog published an article arguing small teams beat tiger teams [4]. The piece self-appoints as president of the "Tiger Team Hater Club™️" and traces the term to NASA's Apollo 13 mission before arguing corporations co-opted it to make "temporary cross-functional committee" sound less boring [4]. The core distinction: tiger teams are temporary specialist units pulled together for a specific problem, while PostHog's small teams are permanent, autonomous, and own full product areas [4]. A second article detailed how PostHog scaled from 1 to 100 IRL events in a year, with ~95% involving customer conversations and over 50% of the company demoing in cities worldwide [5]. The key insight: most dev tool companies struggle to get engineers to demo IRL, but PostHog's "make it public" value and open-source ethos means engineers want to share what they're learning [5]. The piece frames it as "Ship stuff, show people" — the old dev relations adage — but the difference is cultural, not procedural [5].
Brand Humor & Culture
A photo post captioned "What your parents see when you say you're building in public" poked at the gap between open-source ethos and what family members actually perceive [6]. It pulled the highest engagement among authored non-thread posts when retweeted (1,847 views, 63 likes) [7]. A lighter Monday post linked to PostHog's cool-tech-jobs board with "If your Monday's going badly, here:" — a casual employer-branding play [8].
Also this week
Product Announcements & Feature Spotlights (~12%): @posthog announced a Discord AMA with the team behind Replay Vision, with the only off-limits questions being pineapple-on-pizza related [9]. The thread also nudged followers to join the Discord generally — "it's good stuff" [9].
Retweets
1.3k viewsRetweeted @CoastalFuturist congratulating the Tokenmaxxing winner who burned 253.3B tokens in one month, with a link to get a "Member of Token Burning Staff" PostHog hat for non-winners. [10]
123 viewsRetweeted @RafaAudibert celebrating that PostHog's PR #100000 was opened by @veryayskiy (a real person, not a bot) — a milestone that doubles as commentary on how many PRs are now agent-opened. [11]
