A Complete Guide to Using Local AI Models for Everyday Tasks
Most AI tools you’ve used - ChatGPT, Claude, Gemini - run on someone else’s server. You send them your data, they send back an answer, and somewhere a meter is running. For a lot of everyday tasks, that’s overkill. You don’t need a frontier model to write a cover letter draft or sort your reviews into categories, and you definitely don’t need to send that data to a company’s server to do it.
This guide walks through running AI models entirely on your own computer - for free, privately, forever - and picking the right model for the task instead of reaching for the same one every time.
Why bother running models locally?
- Privacy - your data never leaves your machine. This matters more than people think: a resume, a journal entry, a job posting you haven’t told your employer about.
- Cost - once it’s set up, it’s free. No per-request billing, no rate limits, no surprise invoice.
- Availability - works offline, works forever, doesn’t disappear if a company changes its pricing or shuts down a product.
The tradeoff: local models are usually smaller and less capable than the biggest hosted models. For a lot of everyday tasks - drafting text, classifying things, simple Q&A - that gap doesn’t matter.
Two ecosystems, and when to use each
This trips people up, so it’s worth being explicit: there are two different things people mean by “run a model locally,” and they solve different problems.
Ollama - for chat-style, text-generation tasks
Ollama is a tool that runs conversational/text-generation models locally with a simple interface. If the task is “write something,” “answer a question,” “have a back-and-forth conversation” - Ollama is the right tool.
Setup:
1. Download Ollama from ollama.com/download
2. Open a terminal and run: ollama pull gemma3:12b
3. Chat with it directly: ollama run gemma3:12b
That’s it. gemma3:12b is a good general-purpose starting model - capable enough for most writing/summarizing tasks, small enough to run comfortably on a modern laptop.
You can also call it from a script instead of chatting interactively:
import json, urllib.request
def ask_ollama(prompt, model="gemma3:12b"):
data = json.dumps({"model": model, "prompt": prompt, "stream": False}).encode()
req = urllib.request.Request(
"http://localhost:11434/api/generate",
data=data,
headers={"Content-Type": "application/json"}
)
with urllib.request.urlopen(req, timeout=120) as resp:
return json.loads(resp.read())["response"]
print(ask_ollama("Summarize this in one sentence: ..."))
No API key. No account. This is a plain HTTP call to a server running on your own machine.
Hugging Face - for task-specific models
Hugging Face hosts hundreds of thousands of models, but most of them aren’t chat models - they’re built for one specific job: classifying text, detecting sentiment, captioning images, transcribing audio. If your task isn’t “have a conversation” but “label this,” “score this,” “extract this” - you want a Hugging Face model via the transformers Python library, not Ollama.
Finding the right model is the part people get wrong. The homepage sorts by “Trending,” which surfaces huge, bleeding-edge research models that are irrelevant for a simple task. Instead:
- Click the Task you actually need in the sidebar (Text Classification, Image-to-Text, etc.) - not a vague keyword search. Model names don’t always contain the obvious keyword; a well-known sentiment model is literally named after the dataset it was trained on, with no mention of “sentiment” anywhere in its name.
- Drag the Parameters slider down (e.g. under 1B) to filter out giant models you can’t run anyway.
- Sort by Most downloads, not Trending. Downloads are a better trust signal than recency.
- Open the model card and use the
pipeline()code snippet it provides - almost every model page has one.
from transformers import pipeline
classifier = pipeline("sentiment-analysis", model="nlptown/bert-base-multilingual-uncased-sentiment")
print(classifier("This app is great but way too expensive"))
pipeline() handles batches too - just pass a list instead of one string, with a batch_size argument for efficiency. You don’t need to drop into manual PyTorch unless you need very specific control over the model’s internals or you’re operating at a scale where every millisecond matters - neither applies to a personal daily-life tool.
A full worked example
Here’s a real one: a tool that generates a tailored cover letter from your resume and a job posting - entirely offline via Ollama.
Step 1 - write the prompt. The first version looked like this:
prompt = f"""
Their background: {resume_text}
The job posting: {job_text}
Write a cover letter that opens with genuine interest in this specific
role/company (infer company/role name from the posting if present)...
"""
Step 2 - test it, and catch what actually happens. In testing, when no company name existed in the posting, the model didn’t skip mentioning it - it left a literal [Company Name/Role Name - if known, otherwise omit] placeholder bracket sitting in the output. This is a common failure mode: ambiguous instructions get followed literally, not sensibly. Telling a model to “infer if present” doesn’t tell it what to do when it’s not present.
Step 3 - fix the actual ambiguity, not just add more words.
prompt = f"""
Their background: {resume_text}
The job posting: {job_text}
Write a cover letter that opens with a plain greeting ("Dear Hiring
Manager," unless a specific name/company is clearly stated) followed by
genuine interest in the role.
Never use bracketed placeholders like [Company Name] - if you don't know
a detail, just don't mention it.
"""
One sentence removed the ambiguity entirely, and the bug disappeared on re-test. This is the actual rhythm of working with these models: write something reasonable, run it on a real example, look closely at what came back (not just whether it “looks okay”), and fix the specific gap you find - not a vague “make it better” pass.
Step 4 - package it so you’ll actually use it again. A model call buried in a notebook doesn’t get used twice. Wrap it in a small script with clear inputs and outputs:
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("--resume", required=True)
parser.add_argument("--job", required=True)
args = parser.parse_args()
resume_text = open(args.resume).read()
job_text = open(args.job).read()
# ... build prompt, call ask_ollama(), print + save result
Now it’s python3 my_tool.py --resume resume.txt --job job.txt instead of re-copying code into a notebook every time.
Where this actually goes
The same pattern - pick the right ecosystem, find the model with the Tasks filter (not Trending), write a prompt, test it on a real example, fix the specific gap you find, wrap it in a script - applies to almost anything: summarizing your own notes, sorting messages, drafting replies, extracting structured info from a document. The model choice changes; the process doesn’t.
If you want to see this pattern applied end-to-end, PitchCraft - the free tool this example was drawn from - is a working, downloadable version of exactly what’s described above.