Showing posts with label [GIS5103] GIS Programming. Show all posts
Showing posts with label [GIS5103] GIS Programming. Show all posts

Module 6 - Working with Geometries: Extracting Vertex Coordinates from Line Features

The final module of GIS Programming had us write a script that extracts every vertex from a set of river/stream line features (Hawaii stream data) and writes them out to a text file. One line per vertex, containing the feature's OID, a vertex ID, the X/Y coordinates, and the stream's name.

The core of the script is a nested loop: an outer SearchCursor loop iterates over each river feature (row), and an inner loop walks through every point in that feature's geometry (row[1].getPart(0)), incrementing a vertex counter as it goes. For each vertex, I built a formatted string with the OID, vertex ID, coordinates, and stream name, then wrote it to a text file and printed it to the console for verification. 


I hit one bug worth mentioning: I initially tried to pull the river's name using "NAME@" as the cursor field, which failed. The @ suffix in arcpy cursor fields is reserved for special geometry/identity tokens like "OID@" and "SHAPE@." NAME is a normal attribute field, not one of those tokens, so it just needed to be "NAME" on its own. It was a quick fix once I understood why it failed, instead of just matching a pattern I'd seen elsewhere in the script.

One small habit from this module I've kept since: instead of writing the output string directly into both the write() and print() calls, I assigned it to a single variable (txtFormula) first and used that variable in both places. That way, if I needed to change the output format while debugging, I only had to edit one line instead of keeping two copies in sync.



Module 5 - Exploring & Manipulating Data

This module's script had three parts: create a new file geodatabase, copy an existing set of shapefiles into it, then use search cursors to extract and reorganize data from one of those feature classes. First as a filtered printout, then as a populated dictionary.


The first part used CreateFileGDB_management() to create the geodatabase, then looped through every feature class in the source workspace with CopyFeatures_management(), printing a success message and elapsed time for each copy, which is useful for confirming nothing silently failed partway through a batch operation.


The second part used a SearchCursor with an embedded SQL where-clause to pull only the "County Seat" features out of a cities shapefile, printing each city's name, feature type, and circa-2000 population. One thing I noticed working with the data: a few cities show a population of -99999, which is a placeholder for missing/unavailable data rather than an actual value, which is worth flagging if this feature class were used downstream, since a naive average or sum would be badly skewed by it.


This part of the script uses SQL within the SearchCursor to only search for County Seat feature types within the cities attribute table. Using row.getValue() and defined variables, I was able to print each set of data. Below is a summarized flowchart of this section of the script.


The third part is where I hit the most interesting bug of the module. I needed to populate a dictionary (county_seats) using the same cursor logic, but it kept returning an empty dictionary. My first guess was a workspace path issue, which I fixed, but the dictionary still only returned a single entry ({'Carlsbad': 25625}) instead of the full set. I'd assumed that deleting the cursor's current row (del row) was sufficient cleanup between cursor uses, but the real issue was that I hadn't deleted the cursor object itself. Python was holding onto the exhausted cursor from the previous step, so the new loop had nothing left to iterate over. Once I deleted the cursor properly and re-created a fresh one for this step, the dictionary populated correctly with all city entries.


I ended up using the <dictionary_variable>[key] = value method to populate the county_seats dictionary. This final part of the script has a similar flowchart as part B does.

 

I also split one long print statement across multiple concatenated lines (city name / feature type / population) for readability. It was a small style choice, but one that made the script easier to scan when I came back to it later.


 

Module 4 - Geoprocessing

    This module's task was to build a Python script that runs a small spatial analysis pipeline on a hospitals shapefile: add XY coordinates, generate a buffer around each point, and dissolve the resulting buffers into a single feature.

    I structured the script in three stages. First, I used Copy_management() to duplicate the original hospitals shapefile before making any changes as I wanted the source data to stay untouched in case I needed to re-run or compare against it later. Then AddXY_management() added coordinate fields to the copy. Second, arcpy.analysis.Buffer() created a 1000-foot buffer around each hospital point. Third, arcpy.management.Dissolve() merged the buffers into a single output feature. I added arcpy.GetMessages() after each step so I could confirm success (or catch failures) at every stage. The terminal output below shows each step completing with its execution time.

 

    I originally used arcpy.management.Delete() to manually clear old outputs so I could re-run the script during testing. I realized that was the wrong tool as 'arcpy.env.overwriteOutput = True' does this automatically and is the standard way to allow geoprocessing outputs to be overwritten during iterative testing, so I switched to that instead.