It is that time of the year again in the northern hemisphere when parents realize that they forgot to fully plan their summer vacations starting in a few days. And like many others in 2026, I decided to outsource the fixing of this lapse to a frontier AI model.

Twelve months ago, I used an earlier version of the same model to plan a driving trip. A trip for which AI decided to estimate the driving distance as the straight line across fjords. That vacation ended up less of a hiking adventure and more of a Le Mans-style endurance race.

But this time it should have been different. One more year of prompting experience and one more year of agentic AI development by frontier labs.

I asked AI to read all the guides, review all suggestions, and to make a plan for a trip. The model was rather brilliant, hitting all the major highlights in the region. And after I let it know my preferences, it even adjusted the items of interest and helped me navigate the process of finding the tickets. And it recommended a boat ride to visit the river delta. All of this was as I was getting onto the plane.

The night before the boat trip, I asked the model to create a document containing only the single day plan to help me navigate without needing to read the full document. It is only then that AI realized that the trains to the boating docks would not be running due to previously announced maintenance. But not to worry, it told me. Another train line would be available. Except after a bit more prompting to make it check, the model saw that no trains would be available and the trip would have to be the following day.

Now being more on alert, I asked the model to really check everything, including the weather for the newly suggested day. Nothing came back as a concern until I raised the expected near-gale wind gusts myself, and even then I had to explain that wind is one of the reasons river boats get cancelled in the first place. Only at that point did the model agree that this might not be the best time for a leisurely river cruise. And based on the weather chart, which I ended up having to add as a screenshot to convince it to look at more than search summaries, the model recommended the end of the week instead.

A few days later, the night before this new date of the boat trip, I used my super prompt:

"Check all official sites, don't rely on second hand sources, check that there is no maintenance, all locations and stores open, times of the trains, platforms, weather, and where to walk."

I even prompted it to do another paranoid review of all the items and then forced several iterations to make sure it checks for every potential issue. And even uploaded a screenshot of the updated forecast. And the model gave me a confident answer that it really did check everything.

When I got off the boat, it became quite clear that the weather in fact was not fine. There was the small matter of the whole area I was heading to having flooded. AI profusely apologized and then in a panicked register told me that the peak flood is coming a bit later in the day and to try to escape with the first possible boat back. And yes, it then confirmed that even though it knew flooding could be a risk and it was asked to check the weather, it missed the flood warning the forecast agency had issued. It also admitted that this flood risk was clearly there based on the previous day's weather chart.

AI models became far more powerful this year, yet they still cannot safely achieve autonomy on tasks where the only verification available is acting on the model output and observing the potentially unsafe consequences in the real world.

Using a model in such situations means that a human is required to not only review the output of the model, but also both read the sources it relies on and verify that the outputs are complete. Even extensive prompting (in my case, about 25 prompts) was insufficient to catch all the gaps in information. And I was poorly prepared to notice the missing flooding check as I had outsourced the background reading to the AI model.

For tools, agents, and models flooding the market, successful autonomy will require trust. Trust that currently primarily relies on human oversight. And the ability to do this oversight might be eroding with greater reliance on AI.

Looking back, I probably should have finished reading that guide book I got from the library right before leaving. Not only would it warn me about the risk, it would have also given me a bigger appreciation of the cultural history which the model, answering only what I asked, did not volunteer. Which arguably was a bigger miss.