An interactive journalist avatar created in New York — TechCrunch
In New York, the author of a TechCrunch article tested an interactive digital avatar created by Synthesia. The avatar was trained to answer questions only about the journalist's article on why venture capital-backed startups resort to fraud more often than companies without such funding.
To create the avatar, photographs of the author were taken at Synthesia's office and around two minutes of his voice were recorded. The company prepared several versions: personal avatars that voice an entered script, and interactive ones capable of listening to questions and answering them. Each version was created with and without glasses.
How the avatar works
The interactive avatar uses a combination of speech-to-text models, a language model, voice synthesis, and a video model that animates the image while responding. Synthesia uses its own video and voice models, but also allows clients to choose alternative solutions from Cartesia, ElevenLabs, Google, or OpenAI.
More current news is available on the UA.News Telegram channel Telegram.
The journalist's avatar operates in deterministic mode: it responds only within the scope of the article on which it was trained. When asked about his personal life or previous work, it did not answer substantively and redirected the conversation to the article's topic. Creating the avatars took several days.
Synthesia products
Synthesia, originally based in the United Kingdom, develops corporate videos with AI avatars. Last year, the company said its annual recurring revenue had exceeded $100 million, and this year its valuation reached $4 billion.
The company's products include a platform for creating videos from a text script, the Roleplay Sessions service for interactive employee training, and an API platform for combining video and voice models with other services. The article's author noted that the experiment made him consider the use of digital twins in journalism, where trust remains a key element.