Skip to main content

What’s UP Data: The Technical Problems I Solved Building a Geospatial Streamlit App

Building What’s UP Outdoors pushed me beyond notebooks and into geospatial analysis, APIs, interactive mapping, Streamlit state, ETL, and application structure. This post looks at the technical problems I encountered and the Python tools I used to solve them. 

In This Post

My Motivation and Goal

This summer has been full of personal projects to help grow my skills as I work toward my first professional data science role. After attending PyOhio, I was inspired to create What’s UP Outdoors, the biggest project I had taken on since my capstone.

I first started the project about a month ago with heavier AI assistance. After realizing that my understanding was not keeping up with the code, I wrote about the experience and restarted the project from the ground up about two weeks ago. The rebuild gave me a chance to work more deliberately and understand each part of the application as I built it.

What’s UP Outdoors Streamlit map showing the North Country Trail and nearby iNaturalist wildlife observations.
What’s UP Outdoors trail map and nearby nature observations.

I thought the idea would be fairly simple: find hiking trails for my fall-color trip to Michigan’s Upper Peninsula. Instead, the project introduced several technical problems I had never solved before.

Building it pushed me into geospatial data, coordinate reference systems, APIs, interactive maps, Streamlit state, ETL, and application structure.

In this post, I want to break down some of those problems and what I learned while solving them.

How Do I Measure Distance to a Trail?

This was where I started learning GeoPandas, an extension of pandas designed for working with geospatial data. A GeoDataFrame works similarly to a pandas DataFrame, but it also includes a geometry column that can contain spatial objects such as points, lines, and polygons.

One of the most important concepts I had to understand was the coordinate reference system, or CRS.

My ZIP-code coordinates start in WGS 84 (EPSG:4326), which represents locations using latitude and longitude. That works well for storing geographic coordinates and displaying web maps, but it is not the CRS I wanted for calculating distances in Michigan.

For distance calculations, I reproject the data into Michigan GeoRef (EPSG:3078), a projected coordinate system where the coordinates are measured in meters.

def project_point(point):
    """Project a WGS 84 point to Michigan GeoRef."""
    point_series = gpd.GeoSeries([point], crs="EPSG:4326")
    return point_series.to_crs(MICHIGAN_GEOREF).iloc[0]

def distance_to_trails(trails, point):
    """Calculate point-to-trail distances in miles."""
    trails_projected = trails.to_crs(MICHIGAN_GEOREF)
    point_projected = project_point(point)

    return (
        trails_projected.geometry.distance(point_projected)
        / METERS_PER_MILE
    )

These functions project the ZIP-code point and trail geometries into the appropriate CRS, then calculate the distance in miles from that point to each trail’s actual geometry. I use those distances to filter the app so only trails within the user’s selected search radius are shown.

This helped me understand that geospatial locations are more than latitude and longitude values. Trails can be represented as lines or multilines, and those geometries can be used directly in spatial calculations.

That became even more important when I needed to find wildlife observations near an entire trail rather than near a single point.

How Do I Find Wildlife Near an Entire Trail?

After finding trails near a ZIP-code point, I needed to solve a different spatial problem: finding observation points near the full geometry of a trail.

I first create a two-mile buffer around the selected trail. Then I use the rectangular bounding box surrounding that buffer to request nearby observations from the iNaturalist API. Finally, I use GeoPandas to keep only the observations that are actually within two miles of the trail.

def create_trail_buffer(trail, buffer_miles=BUFFER_MILES):
    """Create a buffer around a trail for observation filtering."""
    projected = trail.to_crs(MICHIGAN_GEOREF)
    buffer_meters = buffer_miles * METERS_PER_MILE

    trail_buffer = gpd.GeoDataFrame(
        projected[["TrailGroupName"]].copy(),
        geometry=projected.geometry.buffer(buffer_meters),
        crs=MICHIGAN_GEOREF,
    )

    return trail_buffer

GeoPandas .buffer() creates a polygon around the original trail geometry. Because the trail is represented by a line or multiline geometry, buffering it by two miles creates an area extending two miles around the trail.

As with the earlier distance calculations, I first project the trail into Michigan GeoRef so the buffer distance can be measured in meters.

The iNaturalist API cannot use the GeoPandas buffer polygon directly, so I convert the buffered area into a bounding box the API can use.

west, south, east, north = (
    bounds.to_crs("EPSG:4326").total_bounds
)

params = {
    "swlat": south,
    "swlng": west,
    "nelat": north,
    "nelng": east,
    "d1": start_date.isoformat(),
    "d2": end_date.isoformat(),
    "verifiable": "true",
    "mappable": "true",
    "iconic_taxa": ",".join(TAXON_GROUPS.values()),
    "per_page": PER_PAGE,
    "order_by": "observed_on",
    "order": "desc",
}

After retrieving the API results, I apply a more precise spatial filter.

def filter_observations_near_trail(
    trail,
    observations,
    buffer_miles=BUFFER_MILES,
):
    """Return observations within a specified distance of a trail."""
    distances = distances_to_trail(trail, observations)
    return observations.loc[distances <= buffer_miles].copy()

distances_to_trail() calculates the shortest distance from each observation point to the actual trail geometry, and I keep only observations within two miles.

This taught me an important distinction between retrieval and validation in a spatial workflow. The API bounding box gives me a manageable set of nearby observations to request, while GeoPandas performs the more precise spatial check against the trail itself.

The API does not need to solve the entire problem. It only needs to give my application a useful starting set of data.

How Do I Turn a Data Project Into an App?

Before What’s UP Outdoors, my only experience with Streamlit came from a hackathon. My team built a prediction application, but we divided the work by responsibility. I focused on creating and testing the machine learning model that generated the predictions, while two teammates built the front end with Streamlit.

At the time, I barely thought about the framework itself. I saw Streamlit as something closer to software development and assumed it was outside the main set of tools I needed as a data scientist. Since my part of the project worked without me touching the interface, I never had much reason to investigate it further.

What’s UP Outdoors changed that.

Once I needed a way for users to interact with my trail data, Streamlit made much more sense. It let me turn Python code, pandas and GeoPandas data, API results, and visualizations into an application without building a separate front-end stack.

Getting started was as simple as importing Streamlit with the common st alias and using commands such as st.title(), st.write(), st.caption(), st.image(), and st.subheader().

I could also organize the interface with layout tools such as st.sidebar, st.tabs(), and st.columns().

import streamlit as st

st.set_page_config(
    page_title=TITLE,
    layout="wide",
)

st.sidebar.title("Find Trails Near You")
st.sidebar.write(
    "Enter a ZIP code and choose a search radius to find "
    "nearby Upper Peninsula trails."
)

st.title(f"{TITLE}: Upper Peninsula Trail Explorer")
st.image(BANNER_PATH)
st.subheader(f"How to use {TITLE}:")

tab1, tab2, tab3, tab4, tab5 = st.tabs([
    ":hiking_boot: **Browse & Filter Trails**",
    ":round_pushpin: **Explore Trails Map**",
    ":eagle: **Trail Observations Details**",
    ":sparkles: **AI Trail Summary**",
    ":star: **Saved Favorite Trails**",
])

with tab1:
    st.header("Browse & Filter Trails")

    notes_col, manage_col = st.columns(2)

    with notes_col:
        st.subheader("Trail Notes")

    with manage_col:
        st.subheader("Manage Favorites")

st.tabs() returns a container for each section of the application, and the with blocks determine which Streamlit elements appear inside each tab. I use the same pattern with st.columns() to organize related information side by side instead of putting everything into one long vertical page.

I found the Streamlit documentation easy to follow, along with several YouTube videos that helped me get started. Building the basic interface was surprisingly simple.

The harder part came when I added interaction and realized that Streamlit reruns the Python script when users interact with widgets.

That led to my next problem: how do I keep data and user selections from disappearing or being needlessly recreated after every interaction?

How Do I Keep a Streamlit App from Forgetting Everything?

Streamlit’s execution model and architecture was something I took extra time to understand while building What’s UP Outdoors.

Because Streamlit reruns the Python script as users interact with the app, I needed to think carefully about what should be recalculated, what should be reused, and what should persist for the current user.

I ended up using three Streamlit tools for three different problems:

  • Caching for results that can be reused.
  • Session State for user-specific values that should survive reruns.
  • Fragments for interactions that only need to rerun part of the application.

For example, I cache the processed trail dataset:

@st.cache_data(
    show_spinner="Loading trail data...",
    show_time=True,
)
def load_trails():
    return gpd.read_parquet(PROCESSED_PATH_TRAILS)

trails = load_trails()

st.cache_data prevents Streamlit from repeatedly executing a function when the same inputs have already produced a cached result. I use it for loading the trail and historical observation datasets because those files remain unchanged while someone is using the app.

I also use caching for AI-generated trail summaries:

@st.cache_data(
    ttl="1d",
    max_entries=500,
    show_spinner="Generating Overview...",
    show_time=True,
)
def _generate_trail_summary(trail_data):
    ...

Caching these summaries prevents unnecessary API calls. st.cache_data creates a separate cached result for each unique set of function inputs, so if the trail and observation data have not changed, Streamlit can reuse the previous result.

I set ttl="1d" so each cached summary expires after one day and max_entries=500 so the cache cannot grow indefinitely. When the inputs change, an entry expires, or the cache removes an older entry to make room, Streamlit generates the summary again.

Session State solves a different problem.

if "recent_observations" not in st.session_state:
    st.session_state.recent_observations = {{}}

st.session_state.recent_observations[
    trail_id
] = filtered_api_observations

Some information belongs to the current user’s session rather than to a reusable computation. st.session_state lets those values survive reruns, so I use it for selected trails, favorites, notes, and recently fetched API observations.

Finally, fragments became useful for interactive maps.

@st.fragment()
def build_trail_map(trails, zip_point=None):
    """Build an interactive map of Upper Peninsula trails."""
    map_trails = trails.to_crs(epsg=4326).copy()

A function decorated with @st.fragment can rerun independently, so an interaction inside the fragment does not necessarily require Streamlit to rerun the entire application.

This project pushed me out of thinking only about what my code calculates and into thinking about when it runs.

Learning the difference between caching, Session State, and fragments helped me make the app faster, reduce unnecessary API calls, and prevent user interactions from constantly resetting unrelated parts of the application.

How Do I Put GeoPandas Data on an Interactive Map?

Another fun part of this project was building my first interactive map.

Once my trail geometries were stored in a GeoDataFrame, GeoPandas made basic interactive mapping surprisingly accessible with .explore().

The method creates a Folium map that I can customize with the starting location, zoom level, map tiles, tooltips, and other options.

To display that Folium map inside Streamlit, I use st_folium() from the streamlit-folium package.

trail_map = map_trails.explore(
    location=map_center,
    zoom_start=zoom_start,
    tiles="CartoDB positron",
    tooltip=[
        "HikingName",
        "County",
    ],
)

...

with st.spinner("Loading trail map...", show_time=True):
    trail_map = build_trail_map(
        filtered_trails,
        zip_point
    )

    st_folium(
        trail_map,
        height=600,
        width=1000,
        returned_objects=[]
    )

The location and zoom_start parameters control where the map opens, while tiles controls the background map style. I use tooltip to show trail information when a user hovers over the geometry.

GeoPandas returns a Folium map, which also means I can continue customizing it through Folium when I need more control.

One challenge was that st_folium() can return information about user interactions with the map. I did not need those values for this particular view, so I set returned_objects=[]. That keeps actions such as zooming and panning from sending unnecessary map state back to Streamlit and triggering work I do not need.

Working with GeoPandas and Folium showed me how the same geometry can serve multiple purposes. The trail lines I use for distance calculations and spatial filtering can also become part of an interactive interface.

How Do I Keep the Code Manageable?

This is the most comprehensive and complicated project I have completed so far, so I wanted to keep the code organized as it grew.

In earlier projects, keeping too much logic together made functions harder to find and more difficult to maintain. For What’s UP Outdoors, I wanted each file to have a clear responsibility.

Whats_UP_Outdoors/
├── app.py                         # Streamlit application entry point
├── README.md                      # Project overview and usage
├── SPEC.md                        # Technical MVP specification
├── requirements.txt               # Python dependencies
├── AI_USE_DISCLOSURE.md           # AI-assisted work disclosure
├── static/                        # App images
├── util/                          # Offline ETL workflows
│   ├── etl_dnr_trails.py          # Michigan DNR trail ETL
│   └── etl_inaturalist_history.py # Historical iNaturalist ETL
├── notebooks/                     # Exploratory analysis
├── reports/                       # Generated profiling output
├── data/
│   ├── raw/                       # Local ETL source data
│   └── processed/                 # App-ready Parquet datasets
├── src/                           # Reusable application logic
│   ├── apis/                      # External API helpers
│   │   ├── dnr_api.py
│   │   ├── genai_api.py
│   │   └── inaturalist_api.py
│   ├── ai.py                      # AI summary preparation/generation
│   ├── trails.py                  # Trail processing and filtering
│   ├── locations.py               # ZIP validation and geocoding
│   ├── spatial.py                 # Spatial calculations and filtering
│   ├── inaturalist.py             # Observation processing/summaries
│   ├── streamlit_ui.py            # Streamlit UI helpers and favorites
│   └── maps.py                    # Folium maps
└── tests/                         # Pytest suite

My goal was for app.py to coordinate the workflow rather than contain all of the application logic.

Reusable functions live under src/ and are separated by responsibility. For example, spatial.py handles distance calculations and spatial filtering, maps.py builds Folium maps, and streamlit_ui.py contains reusable interface logic. I also keep API modules separate so external requests remain independent from cleaning, spatial processing, and display logic.

The structure evolved as the project grew. I added streamlit_ui.py after noticing repeated interface logic in app.py, and I moved map-specific work into maps.py for the same reason.

This was also my first time building dedicated extract, transform, and load (ETL) scripts.

In earlier projects, I often cleaned data inside notebooks. That worked well for exploration, but the processing steps were harder to reproduce. Here, the ETL scripts create repeatable workflows that transform the source data into app-ready Parquet files.

I also built the application in small end-to-end slices. Instead of completing an entire layer before moving to the next one, I would take one piece of data from processing through display, verify that it worked, and then add the next feature.

That approach made it easier to validate the data, see whether different parts of the application worked together, and catch problems before I built more code on top of them.

The main lesson I want to carry into future projects is to keep reusable logic modular, make data preparation reproducible, and let the main application file focus on coordinating the workflow.

For the complete project structure and code, see the What’s UP Outdoors GitHub repository.

What Will I Build Next?

Next on my list is returning to machine learning with a new capstone-style project focused on predicting hospital readmissions for diabetic patients.

The topic feels personally meaningful because diabetes has affected people in my family, and it gives me a chance to work through the full data science workflow again. My plan is to move from research and exploratory data analysis into feature engineering, model building, evaluation, and a final presentation.

I also want to revisit my Weather2Go application.

That project gives me a chance to apply what I learned from What’s UP Outdoors, especially around Streamlit, modular code, and reproducible data pipelines. When I revisit Weather2Go, I want to improve the interface, reorganize the project structure, and explore adding a map-based view for weather predictions and their locations.

The two projects will let me build on different parts of what I learned here. Weather2Go will give me another opportunity to strengthen my application development and data engineering habits, while the hospital readmissions project will bring my focus back to machine learning and predictive modeling.

What’s UP Outdoors started as a way to plan a hiking trip. It ended up becoming a much broader lesson in how data moves from raw sources, through spatial analysis and APIs, and into an application someone can actually use.

That is the part of the project I will carry forward most: not just learning new libraries, but learning how the pieces of a larger data project fit together.

Comments