New · Veo 3.1 live — up to 4K with native audio
Intelligence Feed
Guides2026-08-279 min readRichard Byrne

Turn a PDF Into a Video: What Document-to-Video Actually Does in 2026

Wan 3.0 reads a document or a webpage and builds a video from its content. Here is what that genuinely does well, where it falls apart, and what it costs to the credit.

Turn a PDF Into a Video: What Document-to-Video Actually Does in 2026

You can hand a model a PDF and get a finished video back. That is genuinely new in 2026, and it is not what most people assume it means.

I have spent the last week wiring this into our own Studio, which means reading the schema rather than the launch post. This is what document-to-video actually is, what it is good at, and the two places it will waste your money.

What It Is, Precisely

Alibaba's Wan 3.0 accepts a document URL or a public webpage URL and builds a video from what it reads there. Not from a screenshot of the page. From the content.

The distinction that matters: this is not a slideshow generator. It is not converting page one into shot one. The model reads the source, forms an understanding of what it is about, and generates original footage that expresses that. You get a film about your document, not a recording of it.

A single sheet of paper on a dark desk, its printed content rising from the page as a beam of warm light that resolves into moving footage above it

That is why the feature is gated behind the model's reasoning mode. In fal's schema, both file_url and web_url carry the same note: "Requires enable_thinking=true." The model has to comprehend before it can shoot. Our console switches that on automatically whenever you supply a source, because sending the URL without it is accepted and then quietly ignored, which is the worst possible failure: you pay, and you get a generic video that has nothing to do with your document.

A Correction Worth Being Upfront About

When we first covered Wan 3.0, I wrote that document-to-video could not be reached through an API. That was wrong, and the reason is instructive.

I checked which endpoints existed. fal serves three for Wan 3.0: text-to-video, image-to-video, reference-to-video. No document endpoint, so I concluded there was no document input.

Document input is not an endpoint. It is two fields on the reference-to-video endpoint. Checking the endpoint list was simply the wrong altitude to check at, and it produced a confident, published, incorrect sentence that stood for a day.

The lesson is the one this blog keeps earning, sharpened: read the schema, and read all of it. A capability can hide inside an endpoint you already knew about.

Where This Genuinely Earns Its Place

Three cases, from the work we actually do.

A report becomes a trailer. You have a twenty-page industry report. You want ninety seconds that makes someone want to read it. Traditionally that is a scripting job, a storyboard, a shoot or a stock hunt, and an edit. Here the model reads the report and produces footage carrying its themes. You are directing a summary rather than writing one.

A landing page becomes an advert. Point web_url at a live product page and the model reads the positioning, the language, the claims. For a small brand with no existing footage, that is the shortest path from "we have a website" to "we have something to run on Instagram."

A deck becomes a pitch reel. The use case Alibaba led with, and the most obviously commercial. A deck is already a narrative; this turns it into a moving one.

What unites all three: the source is a document of record, and the video is an interpretation of it. That is the job it does well.

Where It Falls Apart

Two failure modes, and both cost real money before you notice.

It will not reproduce your document. If you need the actual numbers, the actual chart, the actual wording on screen, this is the wrong tool and no amount of prompting fixes it. The model is not doing OCR-and-animate; it is reading for meaning and generating fresh imagery. For a video that must show your real Q3 chart, you want a motion-graphics workflow, not a generative one.

Text inside the frame is still unreliable. This is the oldest constraint in generative video and it has not gone away. If your document's key phrase needs to appear legibly on screen, plan to composite it in post. Our own testing on other models found foreground overlay text renders far better than incidental background signage, but "better" is not "dependable" when a client is reading it.

Three nested rectangular frames in warm brass on a dark ground, the innermost holding a small clear image and the outer ones progressively softer, suggesting interpretation rather than reproduction

There is a third limit worth stating plainly: the webpage has to be public. fal's schema is explicit that only pages not requiring a login can be read. Your staging site behind basic auth, your Notion page, your internal wiki: none of those will work, and the failure is not always loud.

What It Costs

Document-to-video runs on Wan 3.0's reference-to-video endpoint, so it is priced exactly like any other Wan 3.0 generation: per second of output, by resolution. There is no document surcharge.

Length720p1080p
10 seconds50 credits100 credits
30 seconds150 credits300 credits

Wan 3.0 Prime, the higher-fidelity tier, runs 40% more: 70 credits for 10 seconds at 720p, 420 for a full 30 seconds at 1080p.

Every figure is on screen in the console before you press Generate, the same as every other model we run. Two practical notes:

Start at 720p. A document-driven generation is an interpretation, and the first one rarely lands. Iterating at 720p costs half what iterating at 1080p does, and the composition you are judging reads perfectly well at that size. Move up when the shot is right.

Use smart duration deliberately, or not at all. Wan 3.0 accepts a null duration, which lets the model choose a length from your prompt and source. That is genuinely useful for a first pass. It is also unpriceable in advance, which is why our console always sends an explicit length: we will not charge you for a duration neither of us chose.

How to Get a Usable Result

The prompt still matters, and this is where most people under-invest. The document tells the model what; your prompt tells it how.

A source alone gives you the model's default interpretation, which tends toward the generic. What changes the output is directing it the way you would direct a DoP who has read the brief:

Cold open on the report's central tension. Handheld, shallow depth of field, desaturated until the resolution lands, then warm. Cut on movement, not on beats.

That is a director's note, not a description, and it is the difference between footage that could illustrate any report and footage that illustrates yours.

If you want the fuller version of that argument, our AI video prompt engineering guide covers the camera and lighting vocabulary these models actually respond to.

Should You Use It?

Use it when the source is the point and the visuals are interpretation: reports, decks, landing pages, articles you want to trail rather than transcribe.

Do not use it when the document's literal content must appear on screen. That is a different craft and generative video is the wrong tool for it, however good the demo looked.

And be honest about the iteration cost. A document-driven generation has more room to misread you than a text prompt does, because there is more input to misread. Budget two or three passes at 720p before you commit to a final render, exactly as you would with any other model.

Common Questions

Can AI turn a PDF into a video? Yes. Wan 3.0 accepts a document URL and generates video from what it reads, rather than animating the pages. It runs in our studio from 50 credits for a 10-second clip at 720p.

Can it turn a webpage into a video? Yes, via a separate web_url field. The page must be publicly reachable with no login, which rules out staging sites and internal wikis.

Will it show my actual charts and numbers? No, and this is the most common misunderstanding. It reads for meaning and generates original footage. If the real chart must appear, composite it in post or use a motion-graphics workflow instead.

What file types work? Alibaba's own platform documents support for doc, xls, ppt, pdf and md. fal's schema simply takes a document URL and states no type restriction, so treat the broader list as the practical guide and test your specific format before committing to it in a client pipeline.

Does it cost more than a normal generation? No. It is the same per-second rate as any Wan 3.0 render, because it runs on the same endpoint.

Try It

Wan 3.0 and Wan 3.0 Prime both run in the Studio today, with document and webpage inputs live, on non-expiring credits. If you would rather hand the brief to someone who does this daily, that is what our production services are for.

document to videoPDF to videoWan 3.0AI videocontent repurposing
Ready to create?

Generate cinematic AI video — from €15

Five frontier models. No subscription. Buy credits, generate on demand, own the results outright.