Dev.to AI 🤖 Ai 👁 0 📖 2 min read

Benchmarking Free AI Models for Outdoor Adventure Generation

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked For Ruta Viva AI, an outdoor exploration app, I wanted to evaluate how reliably free AI models can generate structured outdoor adv

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

For Ruta Viva AI, an outdoor exploration app, I wanted to evaluate how reliably free AI models can generate structured outdoor adventures.

The benchmark measures more than whether a model responds successfully. Each adventure must contain valid JSON, all required fields, exactly three missions, and consistent XP totals.

I designed 30 scenarios across 10 environments, with three prompt variations for each environment. For this initial pilot, I tested three scenarios per model, resulting in nine requests.

Models Tested

I evaluated three free models through OpenRouter:

  • Liquid LFM 2.5 2.6B: liquid/lfm-2.5-2.6b:free
  • NVIDIA Nemotron 3.5 Lightning: nvidia/nemotron-3.5-lightning:free
  • Apodex 1.1 Mini: apodex/apodex-1.1-mini:free

I chose these models to compare structured-output reliability and response speed using a common set of prompts.

Findings

The pilot revealed significant differences in structured-output reliability.

Model Valid JSON Median latency
Liquid LFM 2.5 2.6B 3/3 (100%) 13.57 s
NVIDIA Nemotron 3.5 Lightning 1/3 (33.3%) 41.37 s
Apodex 1.1 Mini 0/3 (0%) 6.36 s

Liquid was the most reliable model in this pilot. All three responses contained valid JSON, the required fields, exactly three missions, and consistent XP totals.

Nemotron produced one valid response; two were truncated at the output limit. Apodex was the fastest by median latency, but none of its responses produced valid JSON, and one returned no usable answer content.

My main takeaway is that speed alone does not determine usefulness. When an application depends on structured data, a fast response that cannot be parsed may be less valuable than a slower, correctly formatted response.

This benchmark has limitations: three requests per model are not enough to establish general performance. The safety check is a basic keyword heuristic, not proof of safety, and adventure quality still requires human evaluation.

Next, I would test more scenarios, repeat requests to measure consistency, and manually evaluate relevance, accessibility, variety, practicality, and safety.

My Benchmark

Public Kaggle notebook: https://www.kaggle.com/code/ezequielsalazar1/outdoor-adventure-ai-benchmark

GitHub repository: https://github.com/Ezequie1Sc/outdoor-adventure-ai-benchmark

The notebook documents the benchmark methodology and validation logic. The results above come from the pilot I ran locally.

kagglechallenge

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.