Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 2 min read

Can an LLM Read a Fraud Network? Testing Gemma 4 31B

This is a submission for the Kaggle Benchmarking Challenge. What I Benchmarked Can a language model use the relationships around a transaction to help identify fraud? I built a benchmark using transaction da

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

Can a language model use the relationships around a transaction to help identify fraud?

I built a benchmark using transaction data from the IEEE-CIS Fraud Detection dataset. Each example presents a target transaction alongside a small text representation of its local network neighborhood. The neighboring transactions are connected through shared attributes such as card, address, or device information. The model does not receive the ground-truth labels in its prompt.

The benchmark measures two behaviors: classifying the target transaction as fraudulent or legitimate, and identifying a potentially suspicious neighboring transaction. I was interested in whether a general-purpose language model could make useful predictions from graph context after that context had been translated into text.

Models Tested

I ran Gemma 4 31B (google/gemma-4-31b) on 100 samples, with both tasks evaluated for each sample. I chose it as the benchmark’s target language model to test whether a large general-purpose model could interpret transaction relationships from the text prompt.

I also compared its classification results with the project’s GraphSAGE graph-neural-network baseline. This was a single-model evaluation, not a broad ranking across language models.

Findings

The run completed with 200 non-empty responses: 100 for transaction classification and 100 for suspicious-neighbor identification. A controlled probe returned text with no token override but an empty response when the 128-token cap was nested under extra_body. The completed run sent the cap as a top-level max_tokens parameter and produced all 200 responses.

Gemma achieved 55% classification accuracy and 53% accuracy on the fraud-neighbor proxy. The 100 classification labels were evenly splitβ€”50 fraudulent and 50 legitimateβ€”so the majority-class baseline was 50%. At 55 correct out of 100, this result does not establish performance above that reference. No baseline comparison is claimed for the fraud-neighbor proxy.

On classification, the GraphSAGE baseline was correct on 67 of the 100 samples. Both systems were correct on 44; Gemma was correct when GraphSAGE was wrong on 11; and GraphSAGE was correct when Gemma was wrong on 23. Both missed 22. In this sample, the graph-based baseline outperformed Gemma.

The neighbor score needs careful interpretation. IEEE-CIS does not provide verified fraud-ring membership or ring-leader labels, so I scored the selected neighbor using its transaction fraud label. That makes 53% a fraud-neighbor proxy score, not evidence that the model can identify actual fraud rings or their leaders.

The main takeaway is that providing a local transaction neighborhood as text did not, by itself, produce strong results from this model on this sample. I would next test more models, increase the sample size, and run prompt ablations to measure how much the neighborhood details affect predictions.

My Benchmark

Kaggle benchmark notebook

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.