Skip to content

Ad Platforms ยท AppLovin

Last Updated: August 11, 2026

Cutdown workflow

How Tierra re-edits longer-form source video (video sales letters, podcast clips, Meta winners over 60s) into AppLovin-native cuts of 60 seconds or less. Read alongside creative for the video specs and the compliance rules.

What a VSL is on AppLovin

A video sales letter is never the interstitial ad itself. It's the long-form source you cut into short AppLovin-native interstitials plus an interactive, and it's often the click destination (the landing page) as well. That's why a long-form pitch is worth cutting: one source feeds a whole batch of short ads and can double as the page they click into. After cutting a VSL, the recommended next step is short 5 to 8 second AppLovin-native, claim-first hooks spliced onto the VSL cuts.

Is it a cut candidate?

Before cutting anything, decide whether the source is worth cutting at all. Judge the concept, not the old cut: is the core hook, angle, and offer compelling on its own? A sound concept is one where the idea holds up and the failure was executional (wrong format, weak edit, poor captions for that platform). A weak concept has no real hook, a flat offer, or nothing that differentiates, and a full 60-second watch won't fix that.

  • Failed on social, concept is sound: a strong AppLovin candidate. Prefer extending or restructuring over trimming, since a long-form pitch can work here where a 15-second scroll-stopper couldn't.
  • Failed on social, concept is weak: skip it.
  • A new source video (one that hasn't run on AppLovin yet) over 60 seconds: run the cutdown workflow. Re-edit to AppLovin structure, don't just trim.
  • A new source already at 60 seconds and working: leave it alone.

Hard specs

Every cut has to meet these (AppLovin's ad specs and requirements):

  • 9:16 portrait, 1080x1920 recommended. Reject 1:1 or 4:5 unless it re-crops cleanly. To take 4:5 or 1:1 source to 9:16 without losing the edges to a crop, center the original frame over a blurred, scaled-up copy of the same footage so the full frame stays visible against a filled background.
  • 60 seconds or less is the documented spec; 59 seconds or less is the defensive target on accounts with length-flag false positives.
  • MP4, H.264, 30fps. Audio AAC 128kbps at -14 LUFS.
  • No embedded call-to-action button. The platform overlays its own, so reserve the bottom one-sixth of the frame for it.
  • Disclaimer (if required): bottom 10%, every frame.
  • The 0 to 6 second hook has to communicate visually with the sound off (about 90% of views are muted).
  • Captions: progressive word highlighting, 2 to 3 words per second max.

Tooling: Whisper plus Gemini

This is one pod's working pipeline, not the only way to run cutdowns; treat it as a proven setup to copy or adapt. In it, two models work together, and Gemini's timestamps are never used for cuts.

Model Role Why
faster-whisper (Whisper) Timestamps, the ground truth Word-level accuracy with word_timestamps=True, vad_filter=True.
Gemini 2.5 Pro Visual notes and editorial signal only Drop its timestamps. They drift 1 to 5 seconds on videos over a minute, verified in a real VSL review cycle. Cut boundaries came out consistently late.

Gemini's output needs a normalizer: its timestamp format drifts from HH:MM:SS.mmm to MM:SS.mmm after the one-minute mark. Strip and re-normalize before anything downstream uses it.

Drive cut selection off the transcript rather than by watching the video. Reading the words instead of the footage raises cut quality and is cheaper. Route only the steps that genuinely need to watch the footage (visual notes, demonstration timing, on-screen action) to a video-capable model like Gemini.

Cost, at Gemini 2.5 Pro pricing (early 2026): about 5 to 10 cents per 4-minute source, with the transcribe stage the larger share. A 20-video batch runs about $1 to $2 of API spend.

Framework timing

Two direct-response frameworks we lean on are Problem-Agitate-Solve and Attention-Interest-Desire-Action. They're proven starting points, not the only options; add others as they prove out on the platform. Whichever you use, target 59 seconds, leaving 1 second of safety.

Problem-Agitate-Solve (the default for direct response):

Beat Timing What
Problem 0 to 15s (0 to 6s is the hook earning the watch) Direct headline, audience callout, pattern interrupt
Agitate 15 to 35s Concrete detail, proof, demonstration
Solution 35 to 55s Mechanism, differentiation
Call-to-action 55 to 59s One or two conversion levers

Attention-Interest-Desire-Action (for awareness or a complex product):

Beat Timing What
Attention 0 to 15s Confident framing
Interest 15 to 35s Storytelling: the situation, the failed solutions
Desire 35 to 55s Mechanism plus differentiation
Action 55 to 59s Strong conversion close

You can't choose placement on AppLovin, so you can't make a video specifically for rewarded versus interstitial. Any impression could be a skippable interstitial or an opt-in rewarded view. Build each cut to work in both: earn attention in the first few seconds in case it's skippable, while still rewarding a full opt-in watch.

Closing call-to-action (final 5s)

Pick one or two conversion levers:

  • Reciprocity or price (discount, free trial, bonus)
  • Urgency (limited time, a deadline)
  • Social proof (customer counts, testimonials)
  • Scarcity (low stock, limited inventory)
  • Clarity (a direct instruction)

Guide the eye toward the bottom of the frame where the platform overlay will sit. Don't compete with the platform's call-to-action; anchor toward it.

Caption rules

AppLovin-native, not a Meta port:

  • Progressive word highlighting (one word at a time becoming prominent), 2 to 3 words per second max.
  • High contrast, legible at glance distance.

Working rule: well-executed captions transfer from other platforms, poorly-executed ones don't. Hard-coded small or low-contrast captions on a Meta video won't work on AppLovin, so re-do them.

Workflow phases

  1. Transcribe (Whisper): ground-truth word timestamps with the VAD filter.
  2. Analyze (Gemini): visual notes, editorial signal, narrative beats. Drop Gemini's timestamps.
  3. Build the brief: a Markdown and JSON brief with cut boundaries from the Whisper timestamps, beat assignments from the Gemini notes, the framework choice (Problem-Agitate-Solve or Attention-Interest-Desire-Action), and the closing call-to-action lever.
  4. Deliver to the editor: the editor cuts to the Whisper timestamps. The Gemini notes are guidance, not law.

Running render and upload at scale

  • Run render and upload jobs on an operating-system-level scheduler, not inside an open application session. Jobs tied to an app that closes die with it, which is a real cause of overnight batch stalls.
  • Make the pipeline idempotent: a re-run skips work already finished instead of redoing it, so a mid-batch failure resumes cleanly.
  • Upload continuously while rendering, rather than batching all uploads at the end, so AppLovin moderation starts early and any rejections backfill within the same wave.

When to stop adding cuts

Saturation guidance:

  • How many ads you can meaningfully test scales with budget, because each ad needs roughly $500 of spend before it's readable. Testing on the order of 250 ads in 2 to 4 weeks takes a high-spend account, very roughly $5k a day and up (about $100k or more of spend over the period). At a $500-a-day launch budget you're meaningfully testing only a handful of ads a week, so match cut volume to the budget.
  • Per source, 5 deliverables per source video hits a ceiling around 50 sources. Past 50 sources, more source diversity gives diminishing returns compared with running more variants per existing source.
  • Quality-control (QC) review, not rendering, is the real throttle. Project-manager (PM) review runs about 2 to 3 minutes per rough cut, so 250 deliverables is 10 to 12 hours of review. Doubling the source count doubles the queue.
  • The diversity picker gives diminishing returns too: the first ~50 picks consume the best variant-by-concept-by-hook-by-spokesperson combinations.
  • Spread cuts for contiguous coverage across the length tiers rather than clustering only at the short end. Gaps in the mid-length tiers underperform, so fill the range (creative has the tier table).

At mass-launch scale the tradeoff flips. Per-asset human review gets dropped: hundreds of rough cuts ship with caption accuracy as the only acceptance bar, and throughput becomes the real constraint rather than polish. The 2 to 3 minute PM review above is the throttle only while you're still reviewing every asset.

Before cutting more from a source, check:

  1. Has the original cut hit any meaningful spend on AppLovin?
  2. Are there fresh angles or hooks to test, or are the remaining cuts within 10% of variants already shipped?
  3. Is the PM review queue under control?

Three yeses, cut more. Any no, pause cutdowns and let what's live earn signal first.

Cost flow (Sheet-driven pipeline)

Tierra's production-rate enabler: the PM drops a Drive URL in column A, runs the Apps Script menu Axon Cuts > Process pending, the backend (Cloud Run, ngrok, or Vercel) produces the Markdown and JSON brief, and the editor delivers an AppLovin-spec 1080x1920 H.264 vertical cut of 60 seconds.

API spend at scale is about $1 to $2 per 20-video batch. PM time is the binding constraint, not API cost, so optimize for PM review throughput (clear briefs, clear cut boundaries, clear beat assignments) rather than cost per cut.