Gold Solutions Partner

Text-to-speech in publisher apps — from a nice-to-have to habit-builder

Posted

For years, audio in publishing was almost synonymous with podcasts, a resource-intensive format that was mostly consumed off platform. Occasionally, a few publishers experimented with narrated articles or long reads voiced by journalists and actors, but these remained niche.

Fast forward to today and AI-driven text-to-speech is reshaping that landscape. The technology has evolved rapidly, offering a scalable way to transform written journalism into a fluid, mobile listening experience. It allows readers to stay with a story, even when they can’t keep their eyes on the screen, and gives publishers a new dimension of engagement with relatively low operational weight.

In this article, we examine how forward-thinking publishers are integrating text-to-speech into their apps, what’s powering this increase in audio and how newsrooms can maximize its value as part of a modern content strategy.

Why audio is now on everyone’s roadmap

Across the U.S., many local and regional publishers are looking for ways to deepen digital engagement as traditional discovery channels become less reliable. The old growth engine driven by search and social referrals is weakening amid the rise of “zero-click” experiences and AI summaries that keep audiences on-platform.

That shift is pushing publishers to think more intentionally about formats that strengthen direct relationships with readers, especially in owned environments like apps.

Our 2025 Media App Report demonstrated that building a deeper connection with audiences increasingly depends on using the full capabilities of mobile devices. Apps that offer richer content formats consistently deliver stronger engagement, with more sessions, longer sessions and more minutes per user each month. Audio is only one part of that mix, but when we looked specifically at apps that include audio, users who engage with it spend nearly twice as much time in the app as those who don’t.

Audience expectations are shifting too. As more publishers invest in audio and video, users are getting used to apps as multi-format environments rather than text-only experiences. For many publishers, particularly those looking to grow digital subscriptions or retain migrating print audiences, text-to-speech is increasingly viewed as table stakes for a dynamic mobile experience.

For many local publishers in the United States, this also connects directly to the long-term shift from print to digital subscriptions. As print readers transition to digital products, publishers are looking for ways to replicate the routines and accessibility that print historically offered. Features like text-to-speech can help bridge that gap, allowing readers to “consume the paper” while driving, walking or doing chores, in ways that feel closer to existing habits.

Turning articles into audio is a low-cost way to increase engagement and habit

Over the past year, using AI-powered text-to-speech technology to translate article content into audio has become an increasingly important component in publishers’ apps. The reason is simple. Readers want to stay close to journalism at moments when reading is awkward or impossible.

For local news publishers in particular, this can open up new moments of engagement. Readers can listen to coverage while commuting, walking the dog, doing chores or driving, moments when reading isn’t practical but staying informed still matters.

Text-to-speech also lets publishers repurpose the content they already produce into an engaging audio experience at relatively low cost. It extends the utility and reach of existing reporting and opens up new times in the day when screens aren’t practical but habit-building is still possible.

It also provides a fast way to trial audio-led experiences and learn what resonates before deciding where human narration or richer production is genuinely worth the investment.

By expanding what a mobile app is usable for, publishers can compete in attention moments they previously struggled to reach. For audiences already accustomed to podcasts or radio, turning articles into audio can increase the utility of a news app.

From a revenue perspective, the first goal is often to increase time spent and frequency of use without asking readers to create more dedicated “sit and read” moments. It can extend sessions in a natural way, especially when playback is smooth, content is queued and the user isn’t forced to constantly navigate.

Commercially, text-to-speech tends to work best as part of a broader value proposition where retention and engagement sit within existing subscriptions, or as a differentiator in premium tiers and membership offers when paired with other benefits like offline access, extra newsletters or events.

How publishers are building listening flows for commutes, catch-ups and long reads

Publishers increasingly see the opportunity to build around moments and formats where reading can require significant effort, like long-form explainers and features that people value but don’t always want to tackle in 10–20 minutes of uninterrupted screen time.

It also works well for “catch-up” flows in the morning or evening that let users stay on top of breaking news while commuting, as well as topic-led listening in areas like sports, business or politics where audiences are happy for the app to keep playing related stories.

Publishers are increasingly leaning into “listen to all” experiences built around a section or edition, alongside thematic playlists. We’re also seeing more hybrid audio editions that combine human-narrated flagship content, such as podcasts, columns and cover stories, with text-to-speech for the wider set of articles.

When publishers package this into something more finite and intentional, like an audio edition or curated playlist, text-to-speech starts to shift from an add-on to a core product feature. At that point, it’s less about scattered one-off plays and more about a structured, time-boxed listening routine that people can return to every day.

In news apps, text-to-speech is often prioritized for live news, explainers and analysis, where “listen while I do something else” moments are most common.

For example, the app from The Independent lets readers listen to “5 things you need to know today” from the top of the home screen, while the The New York Times app features a dedicated Listen tab that blends curated playlists and podcasts with text-to-speech audio articles.

Publishers are mitigating risk with labelling, voices and pronunciation controls

Despite all the interest and clear benefits, text-to-speech isn’t without its challenges.

In an environment where AI is already a source of anxiety for newsrooms, some publishers worry about how synthetic voices will be perceived, especially in hard news, politics and sensitive topics.

As a result, a spectrum of approaches is emerging, from very explicit labelling, of which The New York Times is a notable advocate, to lighter-touch iconography or brief help text.

Underneath that sits a concern that anything blurring the line between human reporting and machine output risks weakening trust that may have taken years to build.

Quality and tone are another major risk. Mispronounced names, places or acronyms can undermine credibility quickly. Tonal missteps on stories involving conflict, tragedy or politics can be just as damaging.

Some of the mitigation is technical and requires teams to correct pronunciations, maintain custom dictionaries or assign different voices to different content types. But much of the perceived risk is editorial and reputational, particularly important in an era where many studies, including from the Reuters Institute for the Study of Journalism, show trust in news declining.

Even as text-to-speech becomes more affordable, it isn’t free, either in direct costs or operational overhead. Publishers need to consider how much audio they need to generate to make this worthwhile, and whether they should roll it out across everything or focus on specific sections and formats.

Product teams also need to prove text-to-speech is adding value rather than simply cannibalizing reading time.

Publishers should generate a clear narrative for how text-to-speech contributes to habit, retention and ARPU. That’s pushing more product and audience teams to treat it not just as a feature, but as an experiment that needs proper measurement.

The next steps

Even over the last year, the voice quality of text-to-speech services has improved dramatically, but there is still a gap between it and a well-produced podcast or strong human narration.

For many users, today’s text-to-speech works well for a quick catch-up or a single article, but it becomes less compelling when you’re asking someone to listen for extended periods. As expectations rise, publishers are increasingly looking for more natural pacing, tighter control over tone and consistent pronunciation for brand-specific jargon, people, places and acronyms.

It’s almost certain the services will continue to improve.

As publishers build more integrated text-to-speech offerings, discovery will matter as much as the listening experience itself. Even the best-designed audio flow will underperform if users don’t encounter it at the right time, so designing a visually appealing UI and educating users through onboarding, tooltips and notifications are important.

The publishers most likely to make the most of text-to-speech will be the ones treating it as a product in its own right, with clear use cases, intentional packaging and measurement.

Text-to-speech won’t be a universal unlock for every publisher. For some it might only be appropriate for a small share of the audience. However, the upside is that, implemented sensibly, text-to-speech is unlikely to be harmful. People who don’t want it will ignore it, and people who do can fold it into their routine, especially when it’s built around defined behaviors.

As consumption continues to shift toward mobile-first, multi-format habits, it’s easy to see why attention on text-to-speech is rising now. Over the coming year, the differentiator will be less about whether a publisher offers text-to-speech at all, and more about who can turn it into something users genuinely rely on.

If you’d like to speak more about your app strategy and how text-to-speech can drive engagement, get in touch with james.kember@pugpig.com. We’ll also be at this year’s Mega-Conference; come and see us at The Pug & Pig pub booth in the Town Square exhibition hall.