Almost every question about unexpected slow motion in generated video has the same answer: you described more action than the seconds you asked for, so the model fitted it in by slowing everything down. This prompt is the case where that is the intended result, and reading it next to that failure mode is instructive.
What makes it safe to ask for is the amount of action. Hovering and a tongue flicking into a bloom. That is it. Two movements, both of them tiny, both of them in one place. Five seconds of that has room to be stretched because there was never fifteen seconds of content fighting for the space.
The lighting is the second condition. Backlit at golden hour with bokeh highlights puts the light behind the subject, which is what makes a wing blur read as a translucent arc rather than as a grey smudge. Slow motion only pays off if the fast thing is visible, and visibility of a fast thing is mostly a lighting problem.
Then there is the air. Pollen and dust in the air is four words and it is doing something disproportionate: it fills the space between the lens and the subject with things that also move slowly, which is how the eye reads the whole frame as slowed rather than just the bird. Without particles, a slow subject against a still background can read simply as a static shot.
One bruised petal is the detail worth stealing regardless of subject. A frame in which everything is perfect reads as rendered. One thing slightly wrong reads as photographed. It is the cheapest single word-count-to-credibility trade in this library.
The sound carries the information the picture is deliberately withholding. Rapid wing hum against garden birdsong and a distant sprinkler: the hum is at its real speed while the picture is not, and that contradiction is exactly what slow motion feels like in a film. A model that writes audio in the same pass can hold both, which a silent model plus a sound pass afterwards would have to be told to do.