Benchmarking Free AI Models for Outdoor Adventure Generation
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked For Ruta Viva AI, an outdoor exploration app, I wanted to evaluate how reliably free AI models can generate structured outdoor adv
This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
For Ruta Viva AI, an outdoor exploration app, I wanted to evaluate how reliably free AI models can generate structured outdoor adventures.
The benchmark measures more than whether a model responds successfully. Each adventure must contain valid JSON, all required fields, exactly three missions, and consistent XP totals.
I designed 30 scenarios across 10 environments, with three prompt variations for each environment. For this initial pilot, I tested three scenarios per model, resulting in nine requests.
Models Tested
I evaluated three free models through OpenRouter:
-
Liquid LFM 2.5 2.6B:
liquid/lfm-2.5-2.6b:free -
NVIDIA Nemotron 3.5 Lightning:
nvidia/nemotron-3.5-lightning:free -
Apodex 1.1 Mini:
apodex/apodex-1.1-mini:free
I chose these models to compare structured-output reliability and response speed using a common set of prompts.
Findings
The pilot revealed significant differences in structured-output reliability.
| Model | Valid JSON | Median latency |
|---|---|---|
| Liquid LFM 2.5 2.6B | 3/3 (100%) | 13.57 s |
| NVIDIA Nemotron 3.5 Lightning | 1/3 (33.3%) | 41.37 s |
| Apodex 1.1 Mini | 0/3 (0%) | 6.36 s |
Liquid was the most reliable model in this pilot. All three responses contained valid JSON, the required fields, exactly three missions, and consistent XP totals.
Nemotron produced one valid response; two were truncated at the output limit. Apodex was the fastest by median latency, but none of its responses produced valid JSON, and one returned no usable answer content.
My main takeaway is that speed alone does not determine usefulness. When an application depends on structured data, a fast response that cannot be parsed may be less valuable than a slower, correctly formatted response.
This benchmark has limitations: three requests per model are not enough to establish general performance. The safety check is a basic keyword heuristic, not proof of safety, and adventure quality still requires human evaluation.
Next, I would test more scenarios, repeat requests to measure consistency, and manually evaluate relevance, accessibility, variety, practicality, and safety.
My Benchmark
Public Kaggle notebook: https://www.kaggle.com/code/ezequielsalazar1/outdoor-adventure-ai-benchmark
GitHub repository: https://github.com/Ezequie1Sc/outdoor-adventure-ai-benchmark
The notebook documents the benchmark methodology and validation logic. The results above come from the pilot I ran locally.
kagglechallenge
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.