Use captions when they already exist
An existing caption track is usually the fastest path and preserves the timing supplied with the source. When no captions are available, Viddash extracts the audio and creates new timed segments with speech recognition.