A recent study led by Matthew Weber and the Computational Research, Media and Organizations Lab at Rutgers University in collaboration with the Local News Impact Consortium (LNIC) found that presence of a local news outlet does not reliably translate into consistent coverage of the surrounding community. From Outlets to Coverage: Understanding New Jersey’s Local News Landscape examined more than 65,000 news articles from 724 local media outlets in New Jersey takes a new approach to measuring local news ecosystems, assessing not just where outlets are located but what they’re actually covering.
To explore the methods behind the study, we asked Weber five questions about what a computational approach to measuring local news ecosystems makes possible and where the field might go from here.
For a long time, the field has measured local news health primarily by counting outlets and weighting them against population. What’s missing when we only look at it that way?
Counting outlets is a useful way to understand the presence of media organizations in a given region, and weighting for population helps to capture an accurate representation of the “density” of media outlets. There are, however, a number of challenges with measuring local news in this way. First, not all media outlets are equal. For instance, in addition to the report on New Jersey’s media ecosystem, the LNIC recently produced a report that looked at story coverage in Montana to better understand the emergence of news deserts in that state, and Valley County, Montana, is a good example of this.
According to the Medill Local News Initiative there is one newspaper in Valley County – the Glasgow Courier. When you actually look at what the Courier produces, you’ll see it’s a weekly newspaper and appears to have only a couple of reporters. The city of Glasgow is one square mile, but the county is more than 5,000 square miles with only 7,519 residents. So does the presence of a media outlet in the county mean it’s not a news desert? Understanding the amount of content produced and the topics covered is important to answering this question.
Consider as well the fact that increasingly media outlets are not headquartered in the counties that they cover. For instance, the Boston Globe moved its headquarters to Dorchester, MA, outside of Boston. The Courier Journal in Louisville, KY is actually headquartered in an office park in Jeffersontown, KY. The location of media outlets doesn’t necessarily correspond to the locations that reporters are covering. These are a couple of the reasons that a more sophisticated measure of content produced and places covered may provide a better sense of what defines local news health.
Without getting too technical, how does a computational approach like this actually work? What is it doing that a researcher or team of researchers manually analyzing articles couldn’t do at scale?
Computational approaches allow us to analyze content at a scale that we simply can’t capture by hand. I worked on a project about a decade ago where we examined local news production across a sample of 100 randomly selected communities. In that project, we coded news articles by hand to determine the topics covered. Our funding supported a dozen undergraduate student coders and the work took hundreds of hours. Today, with computational approaches, we’re able to crawl and analyze websites representing thousands of communities in the same time period.
Using advanced approaches to analyze text, we can now model the topics that are being covered in those articles, and we can also identify places mentioned in the articles to map not only the topics but also the places covered. That can allow for more detail about local media at a far larger scale than traditional methods such as manually coding, where you might check by hand if a place is mentioned in content or not. In addition, in some cases, the speed of advanced methods might make it possible to produce analyses that are more timely or current than those produced using traditional methods. We’ve tested various approaches and we can comfortably crawl smaller sets of websites and produce analyses within a few days.Manual analysis can be slow and incredibly time consuming. The potential advantages of computational approaches are significant and have the potential to reshape what we know about local media ecosystems.
What tradeoffs come with this approach? What does it capture that traditional methods miss, and what might it miss in turn?
At scale, we lose some nuance with computational approaches. For instance, we coded stories based on the critical information needs that they covered. If we have a story that covers a transportation issue but also talks about local politics, the model does not necessarily do a good job deciding whether the primary topic is politics or transportation. But we have some ways to account for this nuance. We use human coders to verify stories where there is a high degree of uncertainty, and we also train our models by having humans code a subset by hand. We aim for a high degree of accuracy – generally around 95%, and these approaches help to improve general accuracy, but there are certainly going to be cases where a story or two is mislabeled or a place or two is miscoded. While there are some limitations, and scale, scope and speed of computational approaches represents a significant gain for researchers.
What would it take for this approach to become something other researchers or funders could use elsewhere?
A core goal of our work is to make these tools available to larger communities of researchers. First, all of our code is available on GitHub and open source. This means that others can use our code and reproduce our methods. Second, we are hoping over the next year to build a user interface that would allow you to run your own analysis without having to know too much about the technical code. Ultimately, our goal is to make these tools accessible to any researcher or practitioner working on local news, not just those with a computational background.
Stepping back, what do you hope this study changes about how the field understands and measures local news impact?
With this study we’re really measuring production. Impact is a next step in this work. Once we know how much content is being produced, what places are being covered, and what topics are being reported on, we need to align that data with what consumers of news are asking for and what they are trying to access. With our Economic Impact Working Group, for instance, we’re hoping to align our data on content coverage with data about community impact and economic impact. Our hope is to find ways to better measure how changes in the nature of local news coverage impacts community members over time.

