Abstract
Mobile device location data offer a low-cost alternative for measuring visitation to outdoor recreation sites and are known to correlate with official visitation counts. Less is known about whether these data can recover recreation demand and consumer surplus comparable to survey-based methods. We compare travel cost models estimated using mobile device and survey data for 17 US National Park Service sites. Results are mixed. We examine the roles of aggregation and sampling bias and use a LASSO regression to assess whether sample and site characteristics explain discrepancies.
1. Introduction
Mobile device location data have become widely available since 2020 and are being used to study many aspects of human behavior. Environmental economists have demonstrated the potential use of these data to study recreation demand, which has long relied on expensive and time-consuming survey methods (Kubo et al. 2020). Mobile device location data offer a relatively low-cost source of high-resolution visitation data across broad geographies over time. Credible welfare estimates over time for a wide range of recreation sites could unlock many research questions and reveal important information about the economic value of public land. The critical research question is whether mobile device location data are a reliable substitute for conventional travel cost surveys.
We seek to answer the substitute question by estimating travel cost models for 17 US National Park Service (NPS) sites using both conventional survey data and mobile device data (Figure 1). The NPS recently launched a visitor survey effort that collects detailed information from park visitors to estimate travel cost models. We compared survey- and mobile-based welfare estimates across a range of park sites, from large and well-known scenic parks to small historic sites. We prepared the data using the same or equivalent assumptions to ensure the most appropriate comparisons. At each site, we estimated a series of models across a spectrum of aggregation. We started by using all the individual-specific information from the individual-level survey data then progressively aggregated the data to investigate the consequences of aggregation characteristic of mobile device data. We estimated models with the mobile device data alone and with augmentation, leveraging information from the survey data. Together, comparison across model specifications and sites allowed us to disentangle the effects of data aggregation from other sources of difference between the consumer surplus (CS) estimates.1
Map of National Park Service and Mobile Sites
Note: The national park site abbreviations are defined in Appendix Table A1.
Recreation demand researchers have sought alternative data sources to better understand recreation values for years. Social media metadata is now stored in central repositories that can be mined to collect data on visitors (Wood et al. 2013; Keeler et al. 2015; Heikinheimo et al. 2017; Sinclair, Ghermandi, and Sheela 2018). Mobility data have become widely available since the late 2010s and contain the information needed to estimate recreation demand models. These data have been used to estimate single-site travel cost models (Kubo et al. 2020) as well as more complex discrete choice models.
Alternative data sources must contain origin and destination visitation data, and the data must accurately describe the behaviors of the population of interest. In general, research has found that social media data can approximate visitation trends to recreation sites (Heikinheimo et al. 2017; Ghermandi 2018; Sinclair et al. 2018; Wood et al. 2020; Wilkins, Wood, and Smith 2021; Goebel et al. 2023). Sessions et al. (2016) show that Flickr data correspond to visitation at US national parks. Using Airsage and Safegraph data, respectively, Merrill et al. (2024) and Tsai et al. (2023) find that mobile data correlate with visitation data at some US national parks but do not correlate well in others. RRC Associates (2024) show how high-resolution data can be used to study how visitors use parks.
Most mobile device location data are now collected via smartphone applications, where users permit location services to collect time-stamped GPS coordinates over time. For most research applications, there is a need to ensure that the sample of devices contributing to the dataset is free from systematic bias. Indeed, Li et al. (2024) compare Advan (our data source) mobile device counts to US Census data and find minimal systematic biases. We conducted similar bias analyses on our sample and draw similar conclusions. We compared the distribution of visit distances (origin to destination) between the survey and mobile device data to investigate systematic differences between the samples of site visitors.
We contribute to this growing literature in three ways. First, we estimate single-site travel cost models for a set of US national parks and derive welfare estimates. Our approach is most similar to Sinclair et al. (2020), who compare travel cost and welfare estimates based on geolocated photographic data and conventional survey data in German national parks. However, we use mobile device data that sample a much larger share of the population and focus on US national park sites that exhibit more variation in site types. To our knowledge, this analysis is one of the first applications of mobile device data to estimate welfare values for recreation sites in the US National Park System. Second, we develop methods to overcome data limitations inherent in mobile device data. Specifically, we trained regression models based on the visitor survey data and use them to impute the probability of flying and the number of people splitting expenses, both of which are key inputs to travel cost models. Third, we compute welfare estimates for NPS sites over time, demonstrating how these data can be used to study values across various timeframes and seasons.
We find that mobile device location data are not a perfect substitute for on-site surveys, in the sense that we cannot simply replace conventional survey data with mobile device data without losing information. However, mobile device data can provide valuable information to fill in temporal and spatial gaps when no surveys can be completed. This can help land management agencies reduce monitoring costs by reducing the required frequency of on-site surveys and has important implications for improved policy and management decisions. For instance, continued comparisons of mobile and survey data can support calibration factors applied to CS estimates in natural resource damage assessments, regulatory analyses, visitor use management strategies, and other analyses that involve recreation opportunities occurring in a season or location with no available survey data. Mobile device data can also play a role in providing information for new legislation. For instance, the 2025 Expanding Public Lands Outdoor Recreation Experiences (EXPLORE) Act (Public Law 118-234) is a significant and wide-reaching piece of bipartisan legislation that seeks to better understand and improve visitor experiences on federal public lands, largely through improved data collection and monitoring.2 Given the considerable resource constraints faced by agencies, supplementing more traditional data-collection methods with mobile device data that are relatively easier to collect over longer time horizons and at dispersed/low-use recreation areas can lead to improved estimates of both recreation visitation and welfare.
2. Data
We assembled data from multiple sources to facilitate a comparison of travel cost models estimated with survey data and mobile device data for a sample of NPS sites across the United States. The survey data are from a new NPS monitoring program that administers visitor surveys at 24 randomly selected park units each year. Mobile device location data were aggregated and provided by Advan using Safegraph Point of Interest locations via DeweyData.io. We supplemented these data with socioeconomic and demographic data from the US Census five-year American Community Survey (ACS). We calculated driving distance and time using Open Source Routing Machine (OSRM).
The NPS launched a visitor survey effort to collect systematic and comprehensive data across a representative sample of park sites. A stratified random sampling approach is used to determine which park sites will be included in the sample each year. First, all park units with adequate visitation data are grouped into one of four categories representing the type of park unit, including natural, recreational, historic urban, and historic nonurban/other. These categories are based primarily on congressional designation and the population of the urban area or metropolitan statistical area in which the park unit is located. Park units are then grouped into a high or low visitation category, based on the top 80% and bottom 20% of park visits in each of the four park unit types. Based on these classifications, every eligible unit of the National Park System is included in one of eight strata. Each year, three parks from each stratified group are randomly selected, resulting in 24 parks surveyed each year. Parks are sampled without replacement to ensure adequate coverage across the system. Our analysis focuses on the first two years of survey administration (2022–2023).
The NPS visitor survey has two components: a relatively short on-site intercept survey that collects information from visitors as tablet-based responses during their visit and a longer follow-up survey that can be completed by the respondent online or returned by mail. The on-site survey is typically administered over a 10-day sampling period and includes key questions for the travel cost modeling, such as the number of trips the respondent took to the park over the past 12 months, the number of people in the respondent’s group splitting trip expenses on the current trip, the transportation modes used by the respondent to travel to the site, and the respondent’s home zip code. To ensure a representative sample of respondents, surveyors are spread throughout each park unit and placed in locations where they are most likely to intercept all types of visitors. Areas such as campgrounds and lodges are typically avoided, since they tend to draw a specific visitor segment. For busier parks, every Xth visitor group is intercepted (X varies depending on how busy the park is), and in each intercepted group, the adult respondent with the most recent birthday is asked to complete the survey. For smaller parks with low visitation, surveyors contact nearly a census of all visitors arriving during the survey period. Although response rates vary across park units, on average, slightly more than 80% of intercepted visitors agree to take the intercept survey. These same respondents are asked to complete the post-trip follow-up survey, which includes additional questions relevant to the travel cost modeling, such as the respondent’s mode of travel and household income. On average, around 30%–45% of visitors respond to the follow-up survey. For more information on the NPS visitor surveys, see Otak Inc., RRC Associates, and University of Montana (2023).
We built a mobile device dataset based on monthly aggregate unique visitors to NPS sites provided by Advan via deweydata.io (Advan Research 2022). Advan constructs monthly aggregate visits to points of interest (POI) based on mobile device GPS time-stamped coordinates from a sample of phones that allow certain apps to use location services. A visit occurs when a device generates a GPS signal from inside a predefined polygon, regardless of dwell time (see the Appendix, under “Mobile Device Data,” for more details on how visits are measured). Consequently, transient devices that pass through a large POI may be counted as a visit. We are unable to change how visits are calculated without access to the underlying data. Moreover, we are unable to change the POI boundary or geofence used to calculate visitation. Some boundaries include the entire site and parking lots, while others contain just an important structure, such as a visitor center. Advan’s monthly patterns product reports these visitation data as different spatial and temporal aggregates. Specifically, they report “visitor_home_aggregation,” the number of unique devices visiting a POI (over a month) by devices’ “home” census tracts in the United States, generating a set of monthly origin-destination pairs necessary to estimate a travel cost model (refer to the Appendix, under “Mobile Device Data,” for how home census tract is determined).
The visit data disaggregated by census tract are subject to “differential privacy” rules, which intentionally distort the data by (1) adding statistical (Laplace) noise to the count, (2) suppressing counts if fewer than two devices visit from a specific census tract (censoring), and (3) reporting four visits if two to four devices visit (truncation). We describe the econometric consequences of the differential privacy in Section 3. Refer to Wan et al. (2026, this issue) for a detailed discussion about differential privacy.
We extracted visits to POIs that best align with the official boundaries of US park sites included in the NPS visitor survey effort. Before December 2022, Advan calculated visitation metrics for all official NPS boundaries. Advan no longer reports visitation metrics based on NPS boundaries and instead reports on specific sites in many US NPS sites.3 In several cases, mobility visitation data existed, but they were too sparse to estimate a travel cost model. We chose a set of 17 sites based on good alignment of park boundaries and sufficient visitation data from both sources. The final list of sites used in our analysis is shown in Appendix Table A1, along with the site location, survey period, and annual visitation.
We aligned the visitation units across the visitor survey and mobile datasets to the greatest extent possible. The NPS survey asks respondents to report the total number of visits taken to the site over the past 12 months. As a result, any repeat visits by the same person in that timeframe are included in the data, regardless of whether those visits occurred in the same week or month. In contrast, the mobile device data report the number of unique visitors by census tract in the calendar month. Multiple visits by the same device within a month are counted once; visits by the same device in different months are counted separately. We aligned the mobile device data with the visitor survey data by summing the monthly visits by census tract over the trailing 12 months before the survey sampling period. Under ideal conditions in which both the NPS survey and mobile device visitors represented the true visiting population, the number of trips reported by survey respondents and the summed unique monthly mobile device visitors by census tract should be comparable.
One of the challenges of using mobility data for travel cost modeling is the lack of individual socioeconomic and demographic information. To investigate the consequences of unobserved data, we estimate survey-based travel cost models using US Census socioeconomic and demographic data, emulating the limitation in the mobility data. We collected both tract and zip code–level data from the 2022 and 2023 five-year ACS on the following variables: total population (Table B01001), median age (B01002), per capita income (B19301), and average household size (B25010). Advan reports visitation by home census tracts using the 2010–2019 official tract boundaries, which do not align with ACS data after 2020. We used relationship files to map the 2022 and 2023 ACS data back to 2010–2019 tract boundaries.4 However, we used the 2010–2019 TIGER geometries to determine origin location in travel distance and time calculations.
Calculating Travel Cost
Travel distances and times are critical inputs to travel cost models. Similar to English et al. (2018), we consider two modes of travel: (1) driving the entire way from home to destination, and (2) flying, which involves driving to the departure airport, flying to a destination airport, and driving to the final destination. For individual i, the weighted travel cost from zip code z to destination j is given by
,1in which dtcijz are the driving-only travel costs, ftcijz are travel costs for both driving and flying, and probflyijz is the probability that someone flies. For NPS survey respondents who reported their mode of transportation, weighted travel costs simplify down to either just the driving-only cost if they drove (because probflyijz in that case is equal to zero) or just the driving and flying cost if they flew (because probflyijz in that case is equal to one). When survey respondents did not answer the question, and for all the mobile visitation data, we impute the probability of flying using a logistic regression model trained on the survey data. Refer to the Appendix, under “Imputation Methods,” for details on the prediction model and results.
We estimated the cost of individual i driving from origin zone z (zip code or census tract) to destination j as
2where cd is the average cost per mile driven, ddjz is the one-way driving distance from the origin to the destination, ch is the cost of a hotel night,
is an integer number of hotel nights assuming one night in a hotel for every 12 hours of driving (which matches the assumption used by English et al. [2018]), nsiz is the number of people splitting the driving cost,5 inciz is the individual’s reported income or the per capita income from zone z, dtjz is the one-way driving time from home zone to destination, and Fj is the per-person fee to access the destination.
We obtained driving distances and times between all survey-based and mobile device origin-destination pairs using OSRM, specifically using the R package “osrm” (Giraud 2022).6 We used polygon centroids of home zip codes (zip code tabulation areas, ZCTAs, for the survey data) and census tracts (for the mobile data) based on US Census TIGER polygons accessed via the R package “tigris” (Walker 2024). We selected popular locations near the entrances of large NPS sites (e.g., visitor centers) as destinations. We acknowledge that the choice of destination can influence travel distance estimations, particularly because large parks can have multiple potential entry points. However, we used the same destination points for the survey and mobile data.
The cd⋅ddjz term of the first sum represents driving costs. We used the American Automobile Association’s (AAA) weighted average operating cost per mile, which considers fuel and maintenance costs per mile averaged over all major vehicle types. The AAA cost estimates were 27.7 cents per mile in 2022 and 25.8 cents per mile in 2023 (AAA 2022, 2023). The chhjz term of the first sum represents hotel costs. For the cost of a hotel night, we used estimates from the American Hotel and Lodging Association, which detail average daily rates for hotels in 2022 ($149) and 2023 ($155) in their 2024 State of the Industry report (Kilic, Cashour, and Carrier 2024).
The second term of the sum represents the opportunity costs of time spent driving. We calculated wage rate by dividing individual- or zone-level income by 2,040 hours (51 40-hour work weeks), which provides an estimate of the number of hours worked per year. We valued time spent traveling by multiplying the wage rate by one-third, as is standard in travel cost literature (Champ, Boyle, and Brown 2017). We multiplied the wage rate by the time spent driving one-way, which is directly provided by OSRM. The first two terms of the sum are multiplied by two, such that costs are on a round-trip basis.
The third and final term in the sum represents NPS site fees. Some of the sites have per-person entry fees, which we used directly for Fj. Other sites have per-vehicle fees, which we divided by the number of people splitting expenses.
Travel costs for those who both drive and fly are more difficult to measure. We decomposed these routes into three legs. First, a person drives from their home to an origin airport. Second, a person flies from an origin airport to a destination airport near the recreation site. Third, a person rents a car and drives from the destination airport to the recreation site. We estimated the cost of individual i flying from zip code z to destination j as
3in which
and
are the one-way driving distances for legs 1 and 3, respectively, rj is the cost of a rental car, aizj is the airline fare from origin airport to destination airport,
and
are the driving times for legs 1 and 3, respectively,
is the flying time for leg 2, and all other terms are defined as in the equation for driving-only travel costs.
We treated an individual’s choice of origin and destination airport as a cost-minimization problem, in which the person considers the 10 nearest origin airports to their home zip code centroid and the 10 nearest destination airports to the recreation site. We narrowed down the list of all US airports to those that appear every year from 2010 to 2022 in the Bureau of Transportation Statistics’ airport rankings for originating domestic passengers (Bureau of Transportation Statistics 2025a). We assigned spatial coordinates to those airports using a Bureau of Transportation Statistics dataset on aviation facilities (Bureau of Transportation Statistics 2025b). This allowed us to identify the closest airports to the home zip code centroids and destination sites, using the R package “sf.” Overall, we assumed that individuals pick the origin-destination airport combo that entails the lowest total travel cost across all three legs (out of the 100 possible combinations across the 10 nearest origin airports and 10 nearest destination airports).
To approximate the cost of a rental car, rj, we obtained average weekly rental car prices at the 15 largest US airports from Nerd Wallet (French 2025). We converted from weekly to daily prices and assigned each US airport a daily rental car price by using the daily rental car price from the closest of the 15 largest US airports. The spatial matching was again done using the R package “sf.” The NPS visitor surveys contain a question asking respondents how many days they planned to spend in the local area near the recreation site. We calculated the median number of local days across all respondents surveyed at a recreation site, which we multiplied by the daily rental car price for all airports near the recreation site.
To calculate airline fares, ajz, we obtained average fares for contiguous US city-pair markets averaging at least 10 passengers per day, as provided by the US Department of Transportation in their Consumer Airfare Report (US Department of Transportation 2025). We assigned every city in the dataset to the NPS region in which it is located, covering the Pacific West, the Intermountain, the Midwest, the Southeast, and the Northeast. We predict average airfare for each city-pair as a function of distance between the cities, year, and region-region combination. We report the linear regression results in the Appendix, under “Imputation Methods.” Using the results, we predict the airline fare for all possible airport-airport combinations.
To obtain time spent flying,
, we divided the geodesic distance between origin and destination (approximated using the R package “sf” package) by the average speed of a commercial passenger aircraft (528 miles per hour).7 Like English et al. (2018), we added two hours to approximate time spent in airports.
Datasets
We constructed a total of three datasets to probe limitations in the mobile device data. First, we built a dataset based on the NPS visitor survey individual trip counts, reported income and demographic information, as well as travel mode and the number of people splitting trip expenses (NPS individual). When respondents did not report travel mode and the number of people splitting expenses, we imputed missing data using the models trained on individual survey data (refer to the Appendix). Second, we limited the NPS survey data further by summing trips by all individuals from each zip code, using five-year ACS aggregate estimates of income and demographics, as well as imputed travel mode and number of people splitting expenses (NPS zonal). These data effectively transform the NPS survey data into a zonal structure similar to the mobility data. The aggregated zonal structure of this data enabled us to study the consequences of aggregation independent of the other sources of variation in the mobile data. Third, we built a dataset based on mobile device visits by census tract, five-year ACS estimates of income and demographics, and imputed travel mode and number of people splitting expenses (mobile). We summarize the datasets in Table 1.
Summary of Datasets Used in Analysis
We constructed the mobile datasets to temporally align with the NPS survey to facilitate the most appropriate comparison of the datasets. We also constructed mobile datasets in each year to investigate changes in recreation demand over time. We used demographics and travel costs associated with 2023, such that the only source of variation is changes in the distribution of travel distances.
Summary Statistics
We present summary statistics in a table and a dot-whisker plot. Trip and travel cost are presented in Table 2 because the variation across NPS sites is not clear in a plot. The demographic variables (age, income, household size) are easier to compare visually (Appendix Figure A1). We focused on the comparison between the NPS zonal and mobile datasets, as the structure and processing of the data are most similar.
Trip and Travel Cost Average by National Park Service (NPS) Site in the NPS Zonal and Mobile Data
Table 2 compares the average trip count per person across the NPS zonal and mobile data. The mean trip counts differ across NPS sites and between data sources. Lake Meredith (LAMR) and Cuyahoga Valley (CUVA) have the highest mean visits in the NPS data, and Great Smoky Mountain (GRSM) and CUVA have the highest mean in the mobile data. The percent difference in visits between the two data sources averages 120%. At nearly all sites, the mobile visits exceed the survey data, suggesting that the mobile data may capture more visitation.
The travel costs show more agreement between the NPS and mobile data sources, with an average percent difference of 67%. Some of the large scenic parks have the highest travel cost in the NPS survey data, including Everglades, Capitol Reef, and Great Sand Dunes (all above $900). Although we similarly find high travel costs for Capitol Reef and Grand Teton in the mobile data, the average travel cost at Everglades is $162. The lower average travel cost to Everglades suggests that the mobile data capture many more local visitors compared with the NPS survey.
Appendix Figure A1 compares the demographic characteristics between the survey and mobile data. We compare aggregate measures of age, income, and household size from the census ACS for both data sources. Differences arise for two reasons. First, the geography differs by data source. The mobile data use census tracts, whereas the NPS data use zip codes. Second, each data source may capture different samples of the population. While there are some differences across sites, the two data sources tend to agree. Age tends to be slightly higher in the mobile data, whereas income is slightly higher in the NPS data.
3. Empirical Model
We specified a series of travel cost models that facilitate the comparison of recreation values between the NPS survey and mobile device data. At its most granular, the NPS survey data contain individual-level trip counts, demographics, and travel cost inputs. However, the mobility data are reported as the aggregate flows from a census tract to a POI. We estimated three types of models corresponding to the datasets described in Section 2: (1) negative binomial travel cost model using individual-level visits from NPS survey data (NPS individual), (2) negative binomial zonal travel cost model using aggregated visits from NPS survey data (NPS zonal), and (3) truncated and censored negative binomial zonal TCM using the mobility data (mobile).
We specified a typical negative binomial TCM for individual-level visits (denoted i) and aggregate visits (denoted z) for each NPS site, j:
4in whichwtcijz is the weighted travel cost,wtcage is age of respondent i or median age of zone z,wtcincome is the income of respondent i or per capita income of zone z, andwtchhsize is the household size of respondent i or median household size of zone z. In the zonal version of the model, we include the population of zone z to absorb variation in trip counts due to differences in population (Hellerstein 1991). Note that we estimate a zonal model using the aggregated data, so the i subscript does not apply. The NPS models are estimated via nbstrat in Stata to accommodate endogenous stratification introduced by the intercept survey (Hilbe and Martinez-Espineira 2005).
We estimated a zonal travel cost model to accommodate the aggregate mobile device visitation data. However, the mobile data are subject to truncation and censoring. The differential privacy applied by Advan does not report data if fewer than two devices visit the NPS site from a single census tract (truncation at one), and reports four visits if two to four devices visit the NPS site from a single census tract (censoring at four). We accounted for this by estimating a negative binomial regression truncated at one and censored at four, with log-likelihood
5in which log(λjz) is the usual log link function with travel cost and demographics defined in equation [4] at the zone level, and yjz is the total number of trips from zone z to site j. f() and F() are the negative binomial probability density function and cumulative distribution function, respectively. The term
accommodates truncation by conditioning the contribution of the zth observation to the log-likelihood function on the probability that an observation greater than one is observed. The term
accommodates censoring by accounting for the probability of observing a two, three, or four when we observe a four in the data. Hypothesis testing is based on robust standard errors.
We calculated CS as
in both the survey and mobile negative binomial models since both use log link functions. Confidence intervals are based on the delta method for the survey models and Krinsky-Robb simulation method in the mobile models (Krinsky and Robb 1986).
4. Results
Our objective was to compare per-trip CS estimates across our subset of parks to determine whether mobility data are a reliable substitute for survey data. In addition, we sought to explain why differences exist between the data sources. We estimated three primary travel cost models for each of the 17 sites for which we have sufficient data; the models converged, and the travel cost coefficients were statistically different from zero. Figure 2 contains the CS estimates and 95% confidence intervals of the three models comparing different levels of aggregation and information: NPS individual labels estimates of the model based on individual-level NPS survey data, NPS zonal labels estimates of the model based on aggregated NPS data, and mobile labels estimates of the model based on the mobile device data (only reported in aggregate form). We include the data underlying Figure 2 in the Appendix.
Consumer Surplus (CS) Per-Trip Estimates from Each Model Used in the Analysis
Note: Mobile = models trained on National Park Service (NPS) data to predict flight probability and number of people splitting expenses; NPS zonal = aggregates trip counts to the zip code and census aggregates for controls; NPS individual = individual information when reported but imputes missing data.
We explored the possibility that the CS estimates based on the survey and mobile data differ because of aggregation. The NPS survey collects individual-level data. We compared CS estimates based on the individual responses (NPS individual) to those from the artificially aggregated NPS data (NPS zonal). We find that the average of the zonal CS estimates ($228) is almost 10% lower than the average of the individual estimates ($252). The Pearson correlation coefficient between the two series of CS estimates is 0.73. The aggregation alone leads to some discrepancy between the CS estimates.
We focused on the comparison between CS estimates from the NPS zonal model and the mobile model (Figure 3). The average of the mobile CS estimates ($214) is about 5% less than the zonal CS estimates ($228). However, the Pearson correlation coefficient is only 0.41, suggesting discrepancies in individual parks. The estimates for Cape Hatteras National Seashore and Great Sand Dunes National Park are within 5% of each other, whereas Tuskegee Airmen National Historic Site and Lake Meredith National Recreation Area differ by nearly 90%. We compared the rank of each NPS site in its respective model. We find that the average absolute difference in park rank between the NPS zonal and mobile models is 4.1 (median 3). Although the majority of sites have similar rank, Gauley River National Recreation Area is ranked number 1 in the NPS zonal model and number 15 in the mobile model.
Plot of Per-Trip Consumer Surplus (CS) Estimates from National Park Service (NPS) and Mobile Device Data
Note: The dotted line is the 45° reference line. Tuskegee Airmen NHS (TUAI) and Gauley River National Recreation Area (GARI) are omitted for clarity. NPS site abbreviations are defined in Appendix Table A1.
Why CS Estimates Differ
We investigated why the CS estimates may differ between the NPS zonal and mobile data. These estimates may differ for several reasons. In large NPS sites with multiple entry points, it may be difficult to sample a representative set of visitors via intercept surveys. The mobile device data are only as accurate as the geofences used to construct the visitation measure. Sites near population centers may erroneously capture foot traffic passing nearby. For example, Gateway Arch National Park includes a walking path over Interstate 44 in St. Louis. A geofence that includes the interstate may simply sample those driving on the interstate. We made sure to use a POI that excludes the overpass. Last, we explored the distribution of visitor travel distances. Discrepancies in the CS estimates, despite similar travel distances, may indicate that assuming census demographics for visitors is inappropriate.
Least Absolute Shrinkage and Selection Operator (LASSO)
We used a LASSO regression to estimate the site and sample characteristics associated with the difference between CS estimates from the survey and mobile datasets. LASSO regressions allow us to include many regressors with very few observations (n = 17) because the penalty will shrink coefficients that do not explain variation (or are redundant to other covariates). We estimated two LASSO regressions using different but related dependent variables: NPS zonal−mobile and the abs(NPS zonal− mobile)/NPS zonal. The set of independent variables includes the following NPS site characteristics: average visits per person, annual visitation (total annual visits five-year average), number of visitor centers, number of campgrounds, large site (binary), high-profile site (binary), historic site (binary), eastern United States, western United States, site total acreage, site total acres of water, percent of local visitors (permanent or seasonal residents of the local area surrounding the park), percentage of visitors whose primary trip purpose was visiting the site, difference in mean travel distance between NPS zonal and mobile data, difference in mean median income between NPS zonal and mobile data, and difference in number of observations between NPS zonal and mobile data. The site characteristics data were obtained from various sources, including the NPS visitor surveys, NPS’s Integrated Resource Management Applications portal, NPS’ Hydrographic and Impairment Statistics database, and an internal NPS Application Programming Interface. All the covariates are scaled (z-score), so the magnitudes are comparable. However, we caution against overinterpretation of the coefficients as LASSO willingly trades bias for better prediction.
We find that NPS site type, size, and nature of visitors correlate with differences between the NPS survey CS estimates and those from the mobile data (Appendix Table A6). The CS estimates from the mobile data exceed the survey-based CS estimates at historic sites (e.g., Mount Rushmore and Tuskegee Airmen), while the opposite is true at sites with a higher percentage of primary purpose visitors. When we seek to explain the absolute value of the difference as a percent of the survey-based CS estimates, we find that larger, more visible parks correlate with smaller differences, while a larger percentage of local visitors and larger income discrepancies correlate with larger CS differences. Together, these results suggest that the CS estimates are more likely to agree at larger parks when there are more sites for visitors to congregate, possibly increasing the intercept rate and measurement at a mobile data POI. On the other hand, the CS estimates are more likely to deviate when there is a larger share of local visitors, indicating that the structure of the aggregate mobile data makes it difficult to measure frequent local visitors. The income discrepancy between the data also points to a difference in who is being sampled.
Travel Distance
We explored the extent to which differences in the distribution of travel distances explain the differences in estimates. Differences in the distribution of travel distances suggest that the survey and mobile datasets capture different visitors and may arise for several reasons. The surveys are generally administered at busy times of the year but may not capture visitors who travel at off-peak times. On the other hand, if the POI geofence does not align with the site boundary, the mobile data may not reflect the true number of visitors.
Appendix Figure A2 depicts the empirical cumulative distribution function (ECDF) of each data source (mobile and NPS) at each site. At most sites, the NPS ECDF lies to the right of the mobile ECDF, suggesting that people traveling farther are more represented in the NPS surveys.8 If travel costs were purely a function of driving distance, this difference in distributions would lead to higher inverse demand curves and thus larger CS estimates. While this difference in travel distances may explain some of the discrepancies in CS estimates (e.g., Great Smoky Mountains National Park), it does not fit in all cases. For instance, the ECDF of Cuyahoga Valley National Park travel distances suggests the NPS sample traveled farther; however, the NPS CS estimates are lower than the mobile estimates. This observation suggests that travel mode and the opportunity cost of time mediate the relationship between travel distance and CS.
Consumer Surplus over Time
One of the advantages of the mobility data is that they are continuously collected over time, allowing us to estimate the zonal travel cost model for any window of time. We estimated models for a subset of NPS sites for 2018–2022 and present the results in Figure 4 (Cape Hatteras National Seashore, CAHA; Grand Teton National Park, GRTE; Mount Rushmore National Memorial, MORU; and Rocky Mountain National Park, ROMO; plots for all other sites are available in the Appendix). Each plot shows the daily CS estimate for visits over the course of the year prior to August 1 of the date listed. We used the 2023 ACS data for all years, so the only source of variation in a park’s travel cost from one year to another is the travel distance and the opportunity cost of time.
Advan Mobility Data Per-Trip Consumer Surplus Estimates by Year for a Subset of National Park Service Sites, 2018–2022
Note: The vertical axes scales differ across sites. The point of interest boundaries used to estimate visitation changed for many parks in December 2022.
The results show different trends in daily CS over time across the four sites. The CS estimates at Cape Hatteras fall in 2020 but rebound to near 2019 levels. Grand Teton estimates peak in 2020 at nearly $330, then fall below $300 by 2022. Other high-profile sites (i.e., Mount Rushmore and Rocky Mountain National Park) show slight increases peaking in 2020 before plateauing or slightly declining. The variation in CS/trip may reflect the strong increase in US national park attendance during the COVID-19 pandemic and the subsequent response to congestion.9 However, the trends we document are not causal evidence of any single explanation.
After the closure during the COVID-19 lockdowns, Rocky Mountain National Park implemented a timed-entry system to limit crowding.10 The system requires visitors to make advanced reservations to enter the park during peak season. It is not clear how the timed-entry system affects per-trip CS a priori. It may reduce visitation, which could decrease CS; however, it may shift the composition of visitors to those traveling from farther away who are more likely to make reservations in advance. We find no evidence that the timed-entry system reduced CS from 2019 to 2022. Further analysis is required to answer this question and is beyond the scope of this article.
5. Discussion and Conclusions
Are mobility data a substitute for travel cost surveys? Our findings suggest that our mobility data are not a perfect substitute for travel cost surveys, despite our efforts to consistently process the data and calculate travel costs. While the CS estimates based on the mobility data capture some heterogeneity across our sample of sites, the mobility estimates do not always align with estimates based on the NPS survey data. Our findings are similar to Tsai et al. (2023), who show that mobility visitation data match well at some NPS sites and not others. We investigate several possible reasons for this discrepancy, including the aggregation of Advan mobility data (along with truncation and censoring), NPS site-specific attributes (i.e., popularity and primary purpose), and sampling and differences in the travel distance distributions. No single factor was found to explain the discrepancies between the survey-based and mobility-based CS estimates.
Implicit in our comparison of results is that the NPS survey data are the gold standard against which we should measure the mobility data. However, NPS surveys are only administered during a few weeks of the year during peak visitation times. This sampling procedure maximizes sample size given a limited budget but misses those who visit outside of that sampling window. If these missed visitors systematically differ by travel distance, income, or travel mode, extrapolating from a sampling window to the entire year may be inaccurate. We show that mobility data can help supplement survey data by providing information about the periods not sampled. This is likely an area where mobile device data can play a critical role. Exploration of this potential use could be beneficial, given the constraints that land managers face in collecting conventional survey data over multiple seasons and time horizons. Perhaps the survey and mobile data could be used in a hybrid-individual-zonal model as proposed in Loomis et al. (2009).
We find that the Advan mobility data are representative of the US population, an important prerequisite of any data purporting to measure visitation. We regress census tract device counts on census population and demographics to assess any systematic biases correlated with observables and find no evidence of large biases. We do find that in census tracts where we observed NPS visitors, device counts were higher where the median age was higher and lower where household size was larger. Higher device counts in tracts with higher median age may reflect the ubiquity of smartphones and overturn the notion that older populations are not represented in mobility data. Lower device counts in areas with larger household sizes are likely a consequence of child privacy laws that prohibit the collection of private data on known underage users.
The literature leveraging mobility data is expanding rapidly, and we build on the area of work focused on understanding outdoor recreation. We contribute to the literature by estimating CS on a sample of US NPS sites using conventional survey data and mobility data. We investigate where the estimates align and where they do not, and we explore several possible underlying reasons. We illustrate how researchers can use survey information to “impute” missing information on mobile device visitors using survey information.
The data in this study are from Advan and are subject to several limitations. First, the data are reported as aggregate visitation metrics by origin census tract as opposed to individual “breadcrumb” data capable of re-creating time diaries. Consequently, we are unable to study multipurpose trips—an important area of the study that alternative mobile data sources are capable of quantifying. Multipurpose trips have always created empirical challenges in the travel cost literature. If a visit to an NPS site is part of a larger bundle of experiences, then fully attributing costs accrued to visit our study site is inappropriate. The aggregated Advan data suffer from the same limitations as other secondary data, such as backcountry permits or camp reservations; we do not observe the entire good. Though the Advan data do contain information on other chain stores visited on the same day as the NPS site, we do not know whether these were incidental visits to a restaurant or destinations. Moreover, we do not observe visits across multiple days. Because we do not account for multipurpose trips in the survey data either, our comparison methods suffer from the same limitations and may overestimate the CS of visitors whose visit is not the sole purpose of the trip.
Advan applies a differential privacy policy that introduces censoring and truncation. We use a negative binomial model that accommodates this censoring and truncation. However, we do not know the consequences of the differential privacy policy on our results. Wan et al. (2026, this issue) study this very issue.
Another limitation of aggregated mobile device data is that aspects of the data processing pipeline are subject to change. In December 2022, Advan discontinued the use of NPS official site boundaries as the geofence to construct visitation metrics and now uses POIs of specific, smaller locations in the park. For instance, Advan now reports visitation on several popular hiking trailheads in Arches National Park, but not the entire park itself. Any changes to the methods of visit attribution affect the fidelity of visit measurement and may create challenges for conducting analyses over time. Advan’s change in geofences effectively cut our visit sample to November 2022, which affected our comparison to NPS data collected in 2023. Recall that we compare the 12 months preceding the survey date. We include an indicator for this affected set of POIs in our LASSO analysis and find no evidence that it explains differences in the CS estimates. In general, researchers using aggregated mobile device data must constantly track changes that data providers make and think through how each change affects their analysis.
Another limitation of the Advan mobility data is the lack of income and demographic information on the visitors. We use census tract aggregates as the best estimate of information about visitors from an origin zone. However, studies have shown that visitors to NPS sites may not be accurately described by aggregate demographics. We emulate this restriction by aggregating the NPS survey data to the zip code level and using ACS values for demographics. We find that the per-trip NPS zonal CS estimates are still reasonably correlated with the NPS individual CS estimates, but they are about 10% less than the individual model estimates. This disparity suggests that estimates from models relying on aggregate data may need calibration but may be capable of tracking changes over time.
Visitation to US national parks remains near an all-time high, suggesting that people derive significant value from this national resource. The NPS depends on reliable estimates of CS to inform a wide range of management decisions. Such estimates can help meet the goals of new legislation, such as the EXPLORE Act, which seeks to improve recreation opportunities on federal lands and waters, largely through improved data collection. Collecting conventional survey data on public lands is costly, time-consuming, and requires complex approval processes. We show that mobility data can supplement traditional travel cost survey methods. In particular, mobility data can highlight the limitations of sampling over a few weeks out of the year and can help identify any systematic biases in those who choose to respond to surveys. However, our findings indicate that the aggregate mobility data from Advan alone, in their current form, are not a reliable substitute for surveys.
Acknowledgments
Bayham thanks Kate Floersheim and Isa Naschold for research assistance. The authors thank Christian Crowley for valuable feedback on an earlier draft. The authors acknowledge the use of ChatGPT, Microsoft 365 Copilot, and Github Copilot for developing code to process and analyze data. However, the written article is entirely the product of the authors. Colorado State University organized, funded, and implemented collection of the mobility data described in this information product. The visitor surveys were organized, funded, and implemented by the NPS. Data were not collected on behalf of the US Geological Survey. Any use of trade, firm, or product names is for descriptive purposes only and does not imply endorsement by the US government. Views and conclusions in this article are those of the authors and do not necessarily reflect the opinions or policies of the NPS.
Footnotes
This article was prepared by a U.S. government employee as part of the employee’s official duties and is in the public domain in the United States.
↵1 We use data collected through the NPS visitor surveys to study the strengths and limitations of mobility data in a pure research context. To facilitate fair comparisons between the survey data and mobile data, we focus primarily on travel cost models that intentionally restrict the survey data (e.g., ignore available information). As a result, the welfare estimates reported here are not necessarily based on travel cost model specifications that result in the most accurate, policy-relevant estimates of CS.
↵2 See https://www.congress.gov/bill/118th-congress/house-bill/6492/text.
↵3 See https://docs.deweydata.io/docs/faqs-advan-research#what-is-the-coverage-of-national-parks.
↵4 See https://www.census.gov/geographies/reference-files/time-series/geo/relationship-files.2020.html.
↵5 We develop a similar strategy to impute the number of people splitting expenses for survey respondents that do not answer the question and for the mobile data. Model details and results are in the ppendix.
↵6 R code is available on Github at https://github.com/jbayham/tc_nps/tree/main and archived at 10.5281/zenodo.18383266. We also tested using the Google Maps API and the ESRI routing server and found the resulting distance and time estimates to be comparable. Unlike OSRM, those alternatives entail a cost. Moreover, the bulk download and caching of driving distance and time is a violation of the Google Maps API terms of service.
↵7 See https://aviex.goflexair.com/flight-school-training-faq/commercial-plane-speeds.
↵8 We estimate Kolmogorov-Smirnov and Anderson-Darling tests to see whether the NPS and mobile samples are from the same distribution. In every case, we reject the null hypothesis that our samples arise from the same unspecified distribution (p-values all less than 0.001). Tests are implemented in R using ks.test() in base R and ad.test() from kSamples (Scholz and Zhu 2023).
↵9 See https://www.nps.gov/subjects/socialscience/visitor-use-statistics-dashboard.htm.
↵10 Refer to https://home.nps.gov/romo/pilot-timed-entry-permit-systems.htm for more information on the timed-entry system.
This open access article is freely available online at: https://le.uwpress.org.










