Large Language Model Selection for Test-Driven Prompt Android iOS Development
TL;DR Highlight
Extended the Python-biased LLM code generation research to Android (Java) / iOS (Swift), and compiled a decision tree for choosing the right model at the right time.
Who Should Read
Android/iOS developers or teams looking to adopt AI code generation in mobile app development. Especially useful when you need criteria for choosing GPT-4o vs open-source models.
Core Mechanics
- 8,704 evaluations on HumanEval & MBPP — directly compared GPT-4o, GPT-4o-mini, Qwen 14B, Qwen 32B across Android (Java) and iOS (Swift)
- TDP (Test-Driven Prompting, a technique that embeds test cases in the prompt to guide correct answers) improved accuracy by an average of +2.22 pp over baseline prompting
- Mobile platform accuracy ranged 66.85%–88.87%, consistently lower than Python code generation (86.90%–91.30%) regardless of model size
- A decision tree for selecting models based on first-attempt accuracy, budget constraints, and self-hosting preference
- Remediation Accuracy (the rate at which incorrect code is fixed on retry) was also measured, providing evaluation closer to real-world workflows
Evidence
- TDP vs baseline prompting: average +2.22 pp (95% CI [1.22–3.23 pp], p < 0.001, Cohen's d = 0.3974)
- Mobile top accuracy 88.87% — similar to or lower than Python's lowest (86.90%)
- 544 programming tasks × 4 models × 2 platforms × 2 strategies = 8,704 evaluations
How to Apply
- Include expected input/output test cases alongside the function signature in your code generation prompt (TDP) to expect roughly a 2 pp accuracy gain.
- If cost matters, consider self-hosting Qwen 32B instead of GPT-4o; attach a Remediation (retry) loop for tasks with low first-attempt accuracy.
- Since mobile code generation quality is lower than Python, build a pipeline that auto-runs unit tests in CI to always validate LLM output.
Code Example
# Test-Driven Prompting Example (Swift function generation)
prompt = """
Write a Swift function that reverses a string.
Requirements:
- Function signature: func reverseString(_ s: String) -> String
Test cases (your output must pass all of these):
- reverseString("hello") == "olleh"
- reverseString("") == ""
- reverseString("a") == "a"
- reverseString("Swift") == "tfiwS"
Return only the function implementation.
"""
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}]
)Terminology
Related Papers
Claude-real-video - any LLM can watch a video
YouTube URL이나 로컬 영상 파일에서 장면 변화 기반으로 핵심 프레임만 추출하고 음성 전사까지 해서 LLM에게 넘겨주는 오픈소스 도구. Claude는 영상 파일을 못 받고, ChatGPT는 자막만 읽고, Gemini는 고정 1fps 샘플링이라는 한계를 모두 우회한다.
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
128K 토큰 컨텍스트에서 모델 내부 attention 신호로 핵심 증거만 추출해 재주입하면 추론 정확도가 24.6% 오른다.
Single and Multi Truth Data Fusion using Large Language Models
여러 소스의 충돌하는 데이터를 GPT-4o-mini 프롬프트로 병합하면 기존 비지도 방법보다 일관되게 F1 점수가 높다.
Multilingual Reasoning Cascades Need More Context
번역 cascade 파이프라인에서 원본 질문을 마지막까지 유지하면 추가 학습 없이 다국어 성능이 크게 오른다.
Less Back-and-Forth: A Comparative Study of Structured Prompting
체크리스트 형식으로 프롬프트를 구조화하면 LLM 답변 품질도 높아지고 토큰도 적게 쓴다.
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
재학습 없이 각 나라의 도덕적 가치관에 맞게 LLM 출력을 조정하는 추론 시점 기법 DISCA 제안
Using Claude Code: The unreasonable effectiveness of HTML
Original Abstract (Expand)
Large language model (LLM) code generation research predominantly focuses on Python, with test-driven prompt engineering exclusively targeting this language. This study presents a comprehensive LLM selection framework for mobile development through rigorous empirical analysis. We conducted 8,704 evaluations across 544 programming tasks (HumanEval and MBPP datasets) on Android (Java) and iOS (Swift) platforms using four state-of-the-art LLMs (GPT-4o, GPT-4o-mini, Qwen 14B, and Qwen 32B), two prompting strategies (base and test-driven), and two metrics (accuracy and remediation accuracy). Systematic analysis of platform-specific patterns yielded a decision tree incorporating first-attempt correctness, budget constraints, and self-hosting requirements, validated through three industry-relevant use cases. Results show test-driven prompting (TDP) achieves a +2.22 pp average accuracy improvement over baseline (95% CI [1.22–3.23 pp], p < 0.001, d = 0.3974). However, LLMs consistently underperform in mobile development (66.85%–88.87%) compared to Pythonbased code generation (86.90%–91.30%) regardless of model size or type. This framework establishes groundwork for platform-specific optimizations while providing practitioners with actionable guidance for model selection in mobile development contexts.