← Back to the site
Oct 2024 – Nov 2024University of Twente

Natural Language Text-to-SQL (LLMs)

Prompt EngineeringNLPLLaMABenchmarking

Anyone who has queried an unfamiliar database knows the problem: you know the question, you just do not know the joins. Large language models can close that gap, and I wanted to know how much of their accuracy comes down to how you phrase the request.

I benchmarked two models, LLaMA-3-SQLCoder-8B and LLaMA-2-7B-chat-hf, on Yale's Spider 1.0 benchmark, putting the same questions to each three ways: as a bare question, as an instruction-led prompt, and as a structured template. Marking it was its own problem, since two queries that look nothing alike can both be right, so I scored exact matches against the reference query and also checked whether actually running the query returned the same rows. I tracked response time alongside accuracy, because a model that is right but slow is no use to someone waiting on a dashboard.

Phrasing helped, but it did not rescue the weaker model. Moving from a bare question to a structured template lifted SQLCoder from 63.3% to 66.9% of queries correct, and chat-hf from 51.3% to 55.5%. That is about three and a half points from prompt design alone on both models, while the twelve-point gap between the two models never closed. Prompt engineering is worth doing, and it is not a substitute for picking the right model.

Accuracy across the three prompt types.
Accuracy across the three prompt types.
← Back to the site