Task Type Model Used VRAM Needed Why
Inline autocomplete, small edits Architecture-level questions ~6 to 8GB Speed matters more than depth here; slow autocomplete gets ignored
Multi-file features, refactors Larger coding-focused model~14 to 16GB Needs more context and reasoning to handle related files correctly
Architecture-level questions Larger model, used deliberately ~14 to 16GB Lower frequency, higher stakes, worth the slower response