Location Bias of my A.I. Assistants - Part III of the Experiments Trilogy
For my Legacy Photo Collection, I decided to select 1,000 photos to upload to my flickr account to share with other photographers, both amateur and professional. I have been taking photographs since I was a teenager, and I purchased my first digital camera in 2003. Therefore, the curation pool logically begins with my 2003 folder and ends with 2025. Even with those start and end dates, I still have about 30,000 photos in my archive.
During the first week or so, I simply scanned through my folders and picked out photographs that appealed to me. I spent about 30 to 60 minutes each day curating roughly a dozen photos, placing them into my Legacy Photo Collection folder and recording them in a spreadsheet with notes explaining why they had been selected.
Prior to my early retirement in 2011, my expeditions typically lasted about a month each. Beginning in 2013, however, they ranged from four to eight months. As the journeys became longer, I found it increasingly difficult to recall every location and the story behind each photograph. I therefore enlisted the help of my AI assistants, believing they would serve as an unbiased second pair of eyes. I was very mistaken.
For a photo like this, most people would not be able to identify where it was taken. I value it because of the story behind it. I know its approximate location because it is stored in my Vietnam folder. My four AI assistants, however, produced dramatically different evaluations, ranging from "discard" to "high overall rating." ChatGPT and Copilot concluded that it did not belong in the Legacy Photo Collection because its technical merit was not above average. Claude noted that the image lacked sharpness but praised its mood, although it could not identify the location. Gemini responded:
"This scene shows the peaceful Lap An Lagoon located in Thua Thien Hue Province, Vietnam. The lagoon is a famous stop for travelers between Hue and Da Nang, surrounded by the Bach Ma mountain range. Small wooden fishing boats, often used for oyster farming, are characteristic of this area."
The actual location was south of Da Nang. In March 2005, I was travelling by minibus from Da Nang to Ho Chi Minh City along the coast. During a 30-minute rest stop, I chose to walk along the beach instead of eating a bowl of pho. A woman wearing a conical hat and carrying a baby on her back approached me, calling out, "One US dollar for a necklace." After a little bargaining, I bought three necklaces for two US dollars. Although I suspected the green "jade" pendants were simply painted stones, I knew the extra money would probably help her family enjoy a better dinner that evening. At the time, I could also have a custom-made cotton dress tailored in Hoi An for only US$10.
Because Gemini was the only AI assistant to identify the correct country, I initially worked with it more frequently during my photo curation. I asked it to locate and rate each photograph after my preliminary selection. The rating system consisted of four criteria: technical merit, artistic value, uniqueness, and travel photo quality. Everything worked smoothly for a day or two. However, after I corrected its location errors, an unexpected pattern emerged. Gemini updated its summaries, but the revised versions often contained mismatched filenames and locations while assigning higher overall ratings. Microsoft Copilot exhibited similar behaviour.
I then wondered whether asking the AI to locate a photograph before rating it was influencing its assessment. What would happen if I reversed the process and asked it to rate the image before identifying the location?
Once again, things worked reasonably well for a few days before Claude encountered similar difficulties. My Croatia folder from my 2017 Marco Polo journey contained 778 photographs taken while I zigzagged through the Balkans. I initially told Claude that my travel route should not matter, but it insisted on knowing it first, so I eventually provided the information. I then selected 32 photographs and asked Claude to identify 26 pics worthy of representing Croatia. Claude's location identifications were mostly incorrect, and together we could verify only 13 of them. Even after updating its summary, 10 filenames were still matched with the wrong locations. After 2.75 hours of back-and-forth discussion, only 10 photographs had both correct ratings and verified locations.
By this point, I wondered whether the problem lay in the subjective rating criteria themselves. I removed the more subjective categories—uniqueness and travel photo quality—and retained only technical merit and artistic value. I suspected that subjective quality assessments, combined with location information, might be causing my AI assistants to reinforce earlier assumptions. To test this idea, I designed an experiment consisting of two parts.
My results supported the hypothesis. At least within this experiment, my AI assistants appeared to exhibit location bias: information about where a photograph was believed to have been taken could influence later evaluations and summaries. The experience left me with an important lesson: avoid providing my AI assistants with route information or background details before asking for an initial qualitative assessment. For now, I have returned to my original workflow—manually screening the photographs and using image searches only when necessary to verify locations. My goal remains the same: selecting 10 to 12 candidates from 300 to 360 photographs in my archive every day.
Comments
Post a Comment