Both of these cost me roughly the same to make. Same setup, same light, same morning. What they cost you later is completely different, and that difference is a procurement question rather than a creative one, which is why it belongs with whoever owns the budget and not only with whoever owns the look.
I shoot both. Here is what each one buys.
The short version
Talking head buys trust you cannot add afterwards. Voiceover buys a file you can change your mind about.
Everything else is detail.
What talking head does that voiceover cannot
A face saying something is a person vouching. The viewer gets tone, hesitation, the small involuntary things that make a claim feel like it belongs to someone. You cannot manufacture that in post and you cannot narrate your way to it.
It is the format for anything where belief is the obstacle: a product that sounds implausible, a category full of overclaiming, a price point that needs justifying. If the viewer's objection is "sure, but does it actually," a face answers that and a voice over B-roll does not. The face has to be one your buyer would believe, though. If that means a different age, skin or hair type from mine, book that creator for the talking head rather than asking my voiceover to carry a job it cannot do.
It is also better for anything with nuance. "It helped, not enormously, but enough that I kept using it" is a sentence that needs a face, because in text or narration it reads as lukewarm and in person it reads as honest.
What voiceover does that talking head cannot
The footage stays editable.
This is the part that gets missed at brief stage, and it is the strongest argument for voiceover by some distance. With voiceover, the visual layer and the language layer are separate files. Which means six weeks later you can change the script, re-record, and ship the same footage saying something different. Different angle, different market, different offer, different season, same shoot.
With talking head, the words are inside the picture. If the messaging changes, the asset is finished. Not degraded, finished. You are back to another booking, another shipment, another fortnight.
For a brand testing seriously, that is a meaningful difference in the working life of the asset. A voiceover video is a platform. A talking head video is a statement.
The practical differences on my end
A few things worth knowing, since they affect what you should ask for.
Voiceover gets more usable footage per hour. I am not delivering lines to camera, so I can shoot the product from more angles and in more setups in the same session. If you want volume, voiceover is the efficient way to buy it.
Talking head is less forgiving of a bad day. Voice can be re-recorded in minutes. A face that is not quite in it is a reshoot.
Voiceover needs better visuals. Nothing is holding attention except the images, so the B-roll has to carry it. Weak footage is exposed immediately, where in a talking head video a face covers a lot.
Sound-off changes the maths. A talking head video with the sound off is a person mouthing at you. Voiceover over clear product action still communicates. Since ads are often watched with the sound off, this matters more for paid than for organic.
How I would choose
Three questions, in order.
Is the obstacle belief or comprehension? Belief wants a face. Comprehension, meaning how it works, what it does and what is in it, is often clearer in voiceover, because the visuals can show while the voice explains.
How long do you need this asset to live? A launch video with a short life can be talking head without regret. Something you intend to run for two quarters is better as voiceover, purely because your messaging will change before the footage stops working.
How certain is the script? If the claims are still moving through approval at your end, do not bake them into a face. A phrase struck after a talking head is shot has no fix: the face already said it. Voiceover survives it — new recording, same pictures.
The hybrid, which is what I usually recommend
Face for the opening, voiceover for the body.
The first few seconds carry the personal vouching that makes a stranger stay, and the opening is doing most of the work anyway. After that, voice over product footage, which keeps the informational middle re-recordable.
You get the trust where it matters and editability where it matters. I shoot both layers in the same session regardless, and I record the voice myself, so a re-record later is a file I send you rather than a session you book. It is what I would pick for most briefs if nobody specified.
What to put in the brief
Not "voiceover please." Tell me how long you want the asset to work for, whether the claims are settled, and whether it is going into paid. I will tell you which of the three this wants, and if we are not sure, say so — I can shoot a talking head version and a voiceover version in one session, which is cheaper than ordering the second one later.
Almost every version of this decision is cheap on the shoot day and expensive afterwards. Make it before I set up, not after you see the file, and then book the session. When the messaging moves, and it will, the second booking should be a voice track rather than a shoot.
Common Questions
Which is better for UGC, voiceover or talking head?
They buy different things. Talking head buys trust that cannot be added in post: a face vouching, with tone and hesitation intact. Voiceover buys an editable asset, because the visual and language layers stay separate files and the script can be re-recorded later against the same footage.
Why does editability matter so much?
Because messaging changes before footage stops working. With voiceover you can ship the same shoot saying something different six weeks later: new angle, market, offer or season. With talking head the words are inside the picture, so a messaging change means another booking, another shipment and another fortnight.
Does the choice change for paid ads?
Yes. Ads are often watched with the sound off, and a talking head video with the sound off is a person mouthing at you. Voiceover over clear product action still communicates, so muted placements tilt the decision towards voiceover.
Is there a middle option?
Face for the opening, voiceover for the body, which is what I would pick for most briefs. Both layers come from one session. The first seconds carry the personal vouching that keeps a stranger watching, and the informational middle stays re-recordable.