Not always. The data requirement depends on the job and the approach. A language feature built on an existing model may need a representative set of requests and approved source material, not millions of training records. A document workflow needs the actual layouts, image quality, handwriting, missing pages, and field variations it will encounter. A prediction system needs enough trustworthy historical outcomes to test whether it improves on the current decision.
The more useful starting point is a small evidence pack: real examples, the result a knowledgeable person would accept, cases the system must reject or escalate, and a record of where the data came from. Someone close to the work must be able to explain why one answer is useful and another is wrong. That domain judgement cannot be recovered from a model catalogue.
If the source data is fragmented, sensitive, poorly labelled, or unavailable through the current systems, that does not automatically end the idea. It changes the first phase. We may need to prove access, establish a baseline, prepare a limited evaluation set, or recommend a non-AI workflow first. The point is to expose the constraint before a larger build depends on it.