Introduction to the Special Issue

Using Cell Phone Mobility Data for Recreation Demand Analysis

Yongjie Ji, Daniel J. Phaneuf and Wendong Zhang

There is a long tradition in environmental economics of using visits to outdoor recreation sites to value environmental goods. Recreation demand (or travel cost) models combine data on trips individuals make to recreation destinations with imputed travel costs and attributes of recreation sites to estimate demand functions for site visits. These functions are used to value access to recreation resources or changes in environmental conditions at recreation destinations. Travel cost studies are ubiquitous, with results supporting benefit-cost analyses, natural resource damage assessments, public lands management decisions, and other policy contexts.

Recreation demand analysis requires information on individual trip-taking, which historically has required primary data collection through costly on-site intercepts or household surveys. This has largely limited the available datasets to single cross sections focused on subnational geographies, describing common recreation activities such as fishing, hiking, boating, and beach-going at well-defined destinations. There are relatively few examples of repeated cross-sectional sampling or tracking households over time, and fewer still examples of national-scale data collection. This has contributed to gaps in the range of places, resource types, activities, and user groups that have been studied. It has also made studying environmental shocks (e.g., oil spills or other sudden changes in site characteristics) difficult due to the inability (without luck or foresight) to collect preshock baseline data. This is especially important in damage assessments, but it also precludes exploiting the types of natural experiments occurring at specific points in space and time that have informed other areas of applied microeconomics.

The recent availability of mobility data offers the possibility of relaxing these data collection constraints. The most common commercial form is app-based: opted-in device users (usually smartphones) share time-stamped GPS locations through their use of applications that record location. Individual place/time stamps are algorithmically aggregated to infer stops of different durations at points of interest during specific timeframes. These passively observed stops, along with information on the approximate home location for a device, can support several distinct analytical data products: origin-to-destination trip flows that link inferred home locations to recreation sites, high-frequency aggregate visitation counts at destinations that enable trend monitoring and event studies, and device-level travel itineraries that reconstruct the full sequence of an individual’s movements. Across these forms, data can be assembled at flexible spatial and temporal scales. Mobility data could therefore be used to study a range of new contexts, such as responsiveness to episodic environmental shocks, emerging threats (e.g., PFAS and wildfires), before-and-after analysis of environmental spills, substitution patterns across large geographies, and diverse activities, among others.

But there are many questions and uncertainties that need to be addressed before the potential of mobility data in recreation demand analysis can be realized. For example, what are the sample properties of datasets constructed using device holders who opt in and out of tracking? What are the quality gradients across different vendors offering various data products, and how does their use of proprietary algorithms affect how we interpret the data? Are there privacy and transparency rules and norms that might limit potential uses? How should we assess validity in general, and relative to traditional approaches? To begin to answer these and other questions, we organized this special issue of Land Economics dedicated to exploring the uses of mobility data in recreation demand modeling.

We shared a call for expressions of interest during late summer 2024 and invited abstracts for consideration during the fall. An online workshop was held in January 2025, at which time authors presented early versions of their research. Manuscripts were submitted and underwent a standard peer review process throughout 2025, with nine articles ultimately accepted for publication in January 2026. We hosted a final workshop with authors, reviewers, and other interested parties in March 2026; the goal was to assemble a list of lessons learned and scope out a continuing research agenda aimed at creating a set of best practices for recreation analysis using mobility data.

The nine articles that follow this introduction provide a wide representation of data types, applications, analysis strategies, vendors, and emphasis points. For example, articles can be categorized based on their use of aggregate origin-destination data provided directly by a vendor or disaggregate device-level stops that were formatted for analysis by the researchers. The geographic scales of analysis include single US states, groups of US states, US nationwide, and an international application. Modeling approaches range from single-site specifications to multisite zonal travel cost models, and from single choice to repeated choice random utility maximization (RUM) models. At least five different mobility data vendors are featured. And while each article addresses challenges specific to its application and data environment, there are several recurring themes. For example, recreation sites are defined as spatial shapes, and trips are recorded when a device stops in a shape. In some cases, analysts use vendor-defined points of interest as the recreation destinations; in others, sites are custom defined by drawing shapes and including buffers around a place of interest. Decisions made on site boundaries affect which stops are considered visits. Many author teams, especially those using aggregate origin-destination data, also needed to address truncation and censoring in their outcome variables due to privacy concerns or the frequency of zeros at high spatial and temporal resolution. All had to wrestle with an imperfect understanding of the sampling characteristics of their data, and several strategies emerged as authors attempted to address representativeness in their context. Some authors benchmark aggregate device visitation counts with administrative data, and some compare their data with demographic distributions from US census data.

We have roughly divided the order of articles based on modeling approach. The first three author teams structure their analyses around the RUM model paradigm. For example, Wan et al. (2026) study the impact of fish consumption advisories on recreational behavior in Michigan in a repeated RUM context. They aggregate device-level data to record recreation visits at the home census block group (CBG)-site-quarter level. This is the unit of analysis in a BLP-style estimator that recovers coefficients on travel cost and advisory status, along with site fixed effects. Cheng and Wan (2026) also use the BLP framework in their study of the welfare costs of the Huntington Beach oil spill. However, assessing the behavioral response in this context requires a shorter time unit, so the authors measure trips at the CBG-site-week level. This introduces the zero-share challenge, so Cheng and Wan (2026) use a recent innovation from the industrial organization literature to augment zero shares using an empirical Bayes approach. Finally, Ahmadiani and Woodward (2026) leverage device-level data from Spectus to reconstruct anglers’ travel trajectories during offshore fishing trips in the Gulf of Mexico, allowing them to model the demand for specific open-water destinations, such as artificial reefs and reefed energy platforms. Together, these articles illustrate both the promise and the technical challenges of applying the RUM model framework to mobility data. All three start from disaggregate device-level stops, though their analytical paths diverge. Wan et al. (2026) and Cheng and Wan (2026) aggregate stops to CBG-site-time matrices, apply BLP-style estimators, and implement empirical Bayes imputation as needed; Ahmadiani and Woodward (2026) retain the device level to infer and reconstruct full travel trajectories. A shared challenge is site definition—whether geofencing water-based parks, handling overlapping beach polygons, or inferring offshore stopping locations through trajectory data—as these decisions determine the effective choice set and impact welfare estimates.

The next three author teams structure their analyses around the trip demand paradigm. Their data sources are also different from the RUM-focused articles: the latter start with device-level data and aggregate for analysis, while these articles start with vendor-provided aggregates. For example, Liu et al. (2026) study the economic value of grassland tourism to Inner Mongolia, China, providing the only non-US application in this special issue. Their cellular signaling data are constructed from cell tower connections, so they are more comprehensive than when location is recorded by application use. The authors observe the weekly flow of visitors from 333 Chinese cities to 102 counties in Inner Mongolia in 2020. They use these data to estimate a zonal travel cost model, illustrating the potential for mobility data to reveal sizable economic value from remote and heretofore understudied ecological resources. Bayham, Enriquez, and Richardson (2026) study the demand for visits to US National Park Service (NPS) sites using single-site demand models. Their emphasis is on assessing the convergent validity of demand estimates based on survey data and mobility data. Specifically, the authors observe monthly device counts to 17 NPS sites, broken out by home census tracts. They also have available visitor survey data collected by NPS at various units, which can be used to estimate conventional travel cost models. The authors find that aggregate cell phone data are not a perfect substitute for individual-level survey data in their context. Finally, Refulio-Coronado et al. (2026) study how water quality at beaches in Rhode Island affects visitation patterns, with particular emphasis on measuring differential response rates between advantaged and disadvantaged communities. Their data provide estimates of monthly aggregate visits from all CBGs in Rhode Island to nearly 800 beach destinations in New England. The authors use a zonal travel cost model to measure how characteristics of the origin CBGs generate heterogeneity in response to site-quality variables, such as beach closures and their duration.

Together, these articles illustrate the versatility of aggregate, vendor-provided origin-destination data across diverse policy contexts—from tourism valuation and equity analysis to monitoring baseline visitation levels relevant to damage assessments—and at substantially lower cost than working with device-level data. A shared challenge, however, is that this aggregation shifts key methodological decisions to the vendor, including privacy-induced censoring that suppresses low-count flows, declining coverage rates since early 2023, and the spatial scope of the recreation destinations. Bayham, Enriquez, and Richardson’s (2026) comparison with NPS visitor surveys raises a validity question that is implicit for all uses of this type of data: how faithfully do aggregate, preprocessed mobility counts represent actual visitation patterns?

The final three articles share an emphasis on investigating how features of mobility data and related modeling choices may impact the validity of estimates. Duff et al. (2026) and Duff, Giguere, and Murray (2026) consider mobility data in the context of natural resource damage assessment (NRDA). In NRDA cases, there is often a premium on measuring the volume of visits to a site before, during, and after an environmental insult. Duff et al. (2026) examine how a range of mobility data products from different vendors performs relative to observed baselines, finding for their case studies that mobility data–derived visit counts do not strongly correspond to directly measured visitation. Duff, Giguere, and Murray (2026) examine the use and limitations of mobility data in a specific NRDA context, showing for a tank fire in Texas that estimates from mobility data display intuitive use patterns before and after the event but may not do a good job of representing the total number of lost trips. These two articles offer a cautionary message on the use of mobility data when the primary objective is to accurately measure aggregate visits to sites rather than the per-trip willingness to pay for a change in a site attribute. Finally, Connolly et al. (2026) use the RUM paradigm in a beach visitation context to investigate the robustness of estimates to several methodological and data-structuring decisions that researchers need to make when using mobility data. They divide decisions into those that are common to all recreation studies and those that are unique to mobility data and systematically compare estimates under different configurations of assumptions. Their findings offer a useful starting point for assembling best-practice guidelines.

Together, these three articles address complementary dimensions of validity: Duff et al. (2026) assess external validity—whether mobility-based visit counts correspond to ground truth—while Connolly et al. (2026) address internal robustness—whether welfare estimates are stable across alternative modeling choices. One early takeaway is that mobility data seem to be better suited to recovering relative behavioral responses and per-trip welfare estimates than to accurately measuring absolute visit volumes.

We believe these articles are an excellent entry point for researchers interested in using mobility data for recreation, and they advance our understanding in important ways. Of course, they are not the final word on this new data resource, and there is still much to learn. There is important research to be done on assessing the validity of mobility data in different contexts and uses, understanding the idiosyncrasies associated with different vendors and data products, and understanding the sampling properties of mobility data and how representativeness compares with traditional, albeit also imperfect, sampling. An important open frontier is the complementary use of mobility and survey data. Beyond being used as benchmarks to help assess representativeness and validity, survey data can be designed to help calibrate and interpret mobility-based estimates, and mobility data can extend the spatial and temporal reach of survey-based analyses. There is also a need for greater transparency from data providers. Researchers would benefit from better documentation of the algorithms used to define stops, the applications contributing location pings, and how privacy-censoring thresholds are set and whether they change over time. There are also new use cases to explore.

Building on discussions from the March 2026 workshop, the author community is developing a lessons-learned document to accompany this special issue, with the goal of establishing a practical reference for best practices in recreation demand analysis using mobility data. We look forward to continuing to engage with this research community.