> ## Documentation Index
> Fetch the complete documentation index at: https://docs.highailabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Testing & Tracking Framework

> How every Professor High video becomes a logged experiment — what we measure, what good looks like, and the rules for scaling a winner or cutting a loser.

## The Premise

We do not know which of these shows will work. Anyone who claims to know which cannabis format goes viral on TikTok is guessing with confidence.

What we can control is **learning rate**. If every video is logged with its script ID, hook ID, show, format, and runtime, then forty videos produce a real answer about what this audience wants. If they are posted without logging, forty videos produce forty opinions.

<Note>
  The rule is simple and absolute: **a video posted without a log row is a wasted test.** The production cost was paid and no information was bought.
</Note>

## What Actually Predicts Reach On TikTok

Not all metrics are equally informative. Ranked by how much they should influence decisions:

<CardGroup cols={2}>
  <Card title="1. Completion rate" icon="circle-check">
    Share of viewers who reach the end. The single strongest driver of distribution. A 30-second video that finishes beats a 60-second video that does not.
  </Card>

  <Card title="2. 3-second hold rate" icon="stopwatch">
    Share still watching at 3 seconds. Isolates hook quality with almost no contamination from the rest of the video.
  </Card>

  <Card title="3. Rewatch rate" icon="rotate-right">
    Views divided by unique viewers. Quiz and reveal formats can exceed 1.0 and that is a very strong signal.
  </Card>

  <Card title="4. Share rate" icon="share-nodes">
    The strongest *human* signal — and the one comedy wins. Laughter is the most shareable emotion. Data earns saves; comedy earns shares.
  </Card>

  <Card title="5. Save rate" icon="bookmark">
    Intent to act later. Our proxy for 'this replaced their budtender.' The metric that best predicts app installs.
  </Card>

  <Card title="6. Comment rate" icon="comment">
    Engagement volume, but noisy — a comment war about indica versus sativa inflates it without indicating quality.
  </Card>
</CardGroup>

<Note>
  **Add share-to-like ratio as an early read.** One brand tracks a **3:1 share-to-like ratio in the first 48 hours** as its success signal; the best-performing mascot video found ran **0.78 shares per like**. It is a sharper 48-hour signal than anything else available, and comedy is where it shows up.

  Rough scale from the teardowns: **above 0.5 shares per like is strong. Above 1.0 is exceptional.**
</Note>

**Follower conversion** (follows per thousand views) is the one that matters for the business but it moves slowly. Read it monthly, not per-video.

<Note>
  **Comedy and data earn different currencies.** Data content earns **saves** — intent to act later. Comedy earns **shares**, because laughter is the most shareable emotion. Shares travel further than saves, which is why the slate is rebalancing toward comedy-led. Judge a comedy episode primarily on share rate and a data episode primarily on save rate; comparing them on the same metric will mislead you.
</Note>

<Warning>
  **Do not optimize for views.** Views are an output, not a lever. Completion rate and hold rate are levers. A video with 400k views and a 22% completion rate taught us less than a video with 9k views and a 71% completion rate.
</Warning>

## Benchmarks

Targets for a new account in its first 90 days. The 3s, 15s, and 30s retention rows are now anchored to [researched platform benchmarks](/social/scripts/platform-research#hook--retention-data) rather than guesses — roughly **70% past 3 seconds**, **60% at 15 seconds**, and **50% at 30 seconds** is the level that signals wider reach.

<Note>
  Around **90% of underperforming videos fail in the opening**, not the body, and the algorithm makes its first distribution call near the **1.5-second mark**. That is why hold rate is weighted so heavily here.
</Note>

| Metric                | Weak       | Acceptable | Strong   | Scale it                     |
| --------------------- | ---------- | ---------- | -------- | ---------------------------- |
| 3s hold rate          | Under 45%  | 45-60%     | 60-70%   | **Above 70%** (reach signal) |
| Completion rate (30s) | Under 25%  | 25-40%     | 40-55%   | Above 55%                    |
| Completion rate (60s) | Under 15%  | 15-28%     | 28-40%   | Above 40%                    |
| Share rate            | Under 0.5% | 0.5-1.2%   | 1.2-2.5% | Above 2.5%                   |
| Save rate             | Under 1%   | 1-3%       | 3-5%     | Above 5%                     |
| Comment rate          | Under 0.3% | 0.3-0.8%   | 0.8-1.5% | Above 1.5%                   |
| Follows per 1k views  | Under 2    | 2-6        | 6-12     | Above 12                     |

<Note>
  Compare 30-second and 60-second videos **only within their own runtime band**. A 60-second video will always lose a completion-rate comparison against a 30-second one, and that comparison means nothing.
</Note>

## The AI-Content Variable

Two researched facts change how results here should be read:

* **Properly labeled AI content carries no algorithmic penalty.** Label every video. Unlabeled synthetic content runs a four-step path to a permanent ban.
* **Low-quality AI loses roughly 30-45% of reach; high-quality AI performs within a few percent of anything else.** So a weak result on a video with visible character drift or artifacts is a *production* failure, not a script failure — do not let it kill a script that was never fairly tested.

Add a `quality_flag` note to the log for any render with visible drift, and exclude those rows when comparing scripts. **Set it at post time** — the first analysis runs at 48 hours, so a flag added later cannot protect it.

## The Test Design

Each batch isolates one variable. Everything else stays fixed. This is the part that is easy to skip and expensive to skip.

| Test               | Variable                              | Held Constant                        | Reads Out After     |
| ------------------ | ------------------------------------- | ------------------------------------ | ------------------- |
| **Hook family**    | Hook ID family (A vs B)               | Same script, same show, same runtime | 6 videos per family |
| **Show format**    | Which show                            | Runtime, posting slot, hook family   | 5 videos per show   |
| **Runtime**        | 30s vs 60s                            | Same show, same hook family          | 8 videos per band   |
| **Comedy vs data** | Register                              | Same show, same runtime              | 10 videos           |
| **Posting slot**   | Time of day                           | Same show, rotating scripts          | 2 weeks             |
| **CTA type**       | Comment-bait vs save-bait vs app link | Same show, same hook                 | 6 videos per CTA    |

<Warning>
  **Never change two variables at once.** A 60-second comedy episode posted at a new time against a 30-second data episode at the old time produces an uninterpretable result and a confident wrong conclusion.
</Warning>

## The First 40 Videos

A concrete plan for the ramp. Four weeks, ten videos a week, and at the end we know which shows to build a season around.

<Steps>
  <Step title="Week 1 — Register and breadth (10 videos)">
    Lead with comedy, since that is where the growth research points. Three from [The Comedy Engine](/social/scripts/comedy-engine), two from Budtender Court, and one each from Receipts, High IQ Test, That's Why You're High, Lab Logs, and Roast My Stash. All 30 seconds. All A hooks. This week answers **comedy vs education** and **which show concept has legs**.
  </Step>

  <Step title="Week 2 — Hook families (10 videos)">
    Run the eight [15s Hook Tests](/social/scripts/hook-tests-15s) plus two re-cuts of week 1's top show with a different hook family. This isolates **hook family effect** and reads out the outcome-led vs claim-led question directly.
  </Step>

  <Step title="Week 3 — Runtime and register (10 videos)">
    Top two shows again. Five at 30 seconds, five at 60 seconds. Within each, split comedy-led against data-led. Answers **how long** and **how funny**.
  </Step>

  <Step title="Week 4 — Double down (10 videos)">
    Eight videos of the winning show, winning hook family, winning runtime. Two wildcards from shows not yet tested, so the slate keeps expanding.
  </Step>
</Steps>

At the end of week 4 there is a defensible answer to: which show, which hook family, which runtime, which register. That is the foundation of a real content calendar.

## The Episode-1 Cliff

Numbered series decay hard. From a 57-video dataset of one animation channel:

| Series   | Ep 1       | Ep 2   | Ep 3  | Ep 4  | Ep 5  |
| -------- | ---------- | ------ | ----- | ----- | ----- |
| Series A | **172.5K** | 114.6K | 60.2K | 67.8K | 63.0K |
| Series B | **240K**   | 222.5K | —     | —     | —     |

**Sequels do not inherit the algorithmic bump.** Combined with comedy having the lowest binge rate of any genre, the implication is direct: **design for a permanent stream of pilots, not seasons.**

Practically, that means every episode must work cold for someone who has never seen the show, and a "part 2" should be a rare deliberate choice rather than a default. When comparing a series' performance, compare **episode 1 against episode 1** of other series — not episode 5 against episode 1, which will always look like failure.

## Decision Rules

Applied at **48 hours** for a first read and **7 days** for the decision. TikTok's long tail is real — a video can find an audience on day 5 — so no video is judged dead before day 7.

```mermaid theme={null}
flowchart TD
    A[Read at 7 days] --> B{3s hold rate?}
    B -->|Strong| C{Completion?}
    B -->|Weak| D{Completion among<br/>those who stayed?}
    C -->|Strong| E["✅ SCALE<br/>4 more, same show,<br/>same hook family"]
    C -->|Weak| F["🔁 ITERATE<br/>keep hook,<br/>rewrite the middle"]
    D -->|Decent| G["✂️ RE-CUT<br/>swap to the B hook.<br/>Do not re-shoot"]
    D -->|Weak| H{Third weak result<br/>for this show?}
    H -->|Yes| I["🛑 CUT<br/>retire the slot,<br/>update Status Board"]
    H -->|No| J[Keep testing]
```

<CardGroup cols={2}>
  <Card title="Scale" icon="arrow-trend-up">
    Any metric in the "Scale it" column, or two in "Strong." Produce four more episodes of that show with that hook family immediately. Winners decay — move while it works.
  </Card>

  <Card title="Iterate" icon="rotate">
    Strong hold rate, weak completion. The hook works and the body does not. Keep the hook, rewrite the middle, shorten by 15 seconds.
  </Card>

  <Card title="Re-cut" icon="scissors">
    Weak hold rate, decent completion among those who stayed. The content is fine and the opening is not. Re-cut with the B hook. Do not re-shoot.
  </Card>

  <Card title="Cut" icon="ban">
    Three consecutive episodes in the "Weak" column across both hold and completion. Mark the show `retired` on the [Status Board](/social/show-status-board) and move the slot. Sunk cost is not an argument.
  </Card>
</CardGroup>

<Note>
  **Repeat what works, immediately and shamelessly.** When an episode outperforms, the next three videos should be near-clones — same show, same hook family, same runtime, different subject. This is the single highest-return behavior in short-form video, and the one most creators resist because it feels repetitive to the maker long before it feels repetitive to the audience.
</Note>

## The Performance Log

The log lives at `social-content/output/tracking/performance-log.csv` in the repository. One row per posted video, filled at post time and updated at 48 hours and 7 days.

| Column                                                    | Filled When | Notes                                                                                                                                                                                         |
| --------------------------------------------------------- | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `script_id`                                               | Post        | e.g. `TR-001`. Joins to the script pack.                                                                                                                                                      |
| `hook_variant`                                            | Post        | `A`, `B`, or `C`.                                                                                                                                                                             |
| `hook_id`                                                 | Post        | e.g. `DD-01`. Joins to the [Hook Library](/social/scripts/hook-library).                                                                                                                      |
| `hook_family`                                             | Post        | `BC`, `KG`, `PA`, `DD`, `ST`, `CT`, `QN`, `CM`, `OL`, `PI`. All ten families in the [Hook Library](/social/scripts/hook-library).                                                             |
| `show`                                                    | Post        | Show slug.                                                                                                                                                                                    |
| `runtime_sec`                                             | Post        | Actual, not intended.                                                                                                                                                                         |
| `register`                                                | Post        | `data`, `comedy`, or `mixed`.                                                                                                                                                                 |
| `cta_type`                                                | Post        | `comment`, `save`, `follow`, `site`, `app`, `share`, or `none`. Matches the [CTA ladder](/social/funnel#the-cta-ladder); `none` is a catchphrase-only close.                                  |
| `platform`                                                | Post        | `tiktok`, `reels`, `shorts`.                                                                                                                                                                  |
| `posted_at`                                               | Post        | ISO 8601, with timezone.                                                                                                                                                                      |
| `post_url`                                                | Post        | Direct link to the social post.                                                                                                                                                               |
| `dest_url`                                                | Post        | **The tracked destination link this video used** — unique short link or UTM-bearing URL per script, never the bare domain. Required for every Traffic-arm script or the arm cannot be judged. |
| `views_48h` / `views_7d`                                  | 48h / 7d    |                                                                                                                                                                                               |
| `hold_3s_pct`                                             | 48h         | From TikTok analytics.                                                                                                                                                                        |
| `completion_pct`                                          | 48h         |                                                                                                                                                                                               |
| `avg_watch_sec`                                           | 48h         |                                                                                                                                                                                               |
| `shares_48h` / `saves_48h` / `comments_48h` / `likes_48h` | 48h         | Separate from the 7d counters — these are cumulative, so sharing a column would overwrite the 48-hour reading and lose the share-to-like signal above.                                        |
| `shares_7d` / `saves_7d` / `comments_7d` / `likes_7d`     | 7d          |                                                                                                                                                                                               |
| `follows`                                                 | 7d          |                                                                                                                                                                                               |
| `site_sessions_7d`                                        | 7d          | Sessions attributed to `dest_url`. The primary metric for the [Traffic arm](/social/scripts/the-100/design).                                                                                  |
| `verdict`                                                 | 7d          | `scale`, `iterate`, `recut`, `cut`.                                                                                                                                                           |
| `quality_flag`                                            | **Post**    | `drift` if the render has visible character drift or artifacts. Set at post time, not day 7 — the first analysis runs at 48h and a blank flag passes every predicate.                         |
| `notes`                                                   | 7d          | One line. What we think happened and why.                                                                                                                                                     |

The `notes` column is the most valuable one in the file and the easiest to leave blank. A number tells us what happened; the note is the only place the *why* survives long enough to be useful three months later.

## Weekly Review

Thirty minutes, once a week. Four questions, in order:

<Steps>
  <Step title="Which hook family won?">
    Average 3s hold rate grouped by `hook_family`. Any family with fewer than four data points is noise, not a result.
  </Step>

  <Step title="Which show won?">
    Average completion rate grouped by `show`, compared within runtime bands.
  </Step>

  <Step title="What is next week's clone batch?">
    The winning combination gets at least six of next week's ten slots.
  </Step>

  <Step title="What gets cut?">
    Apply the cut rule without negotiating with it. Update the Status Board the same day.
  </Step>
</Steps>

## Where This Connects

Individual episode drafts are tracked in the `social_content_drafts` table in Supabase, surfaced through the [Social Content Drafts](/help/features/social-content-drafts) queue. That table tracks **draft workflow state** — generated, reviewed, approved.

This CSV tracks **published performance**. They are deliberately separate: the draft queue answers "what is ready to post," and the log answers "what should we make more of." Joining them on `script_id` is the eventual path to an automated dashboard, once there are enough rows for that to be worth building.

## Related

<CardGroup cols={3}>
  <Card title="Hook Library" icon="bolt" href="/social/scripts/hook-library">
    The hook IDs this framework measures.
  </Card>

  <Card title="Script Room" icon="clapperboard" href="/social/scripts/overview">
    The scripts under test.
  </Card>

  <Card title="Status Board" icon="table-list" href="/social/show-status-board">
    Where scale and cut decisions get recorded.
  </Card>
</CardGroup>
