Tanui Kipng'etich Sila

Computer Scientist

Home About Education Experience Skills Projects Blog Contact
Home / Blog / Artificial Intelligence
Sep 19, 2026
• Artificial Intelligence
• 20 min read

Creating Knowledge-Aware AI Applications with Python and OpenAI API

Tanui Kipng'etich Sila
17 views 0 comments

RAG using Python Prompts

A Practical Step-by-Step Guide to Building Knowledge-Aware AI Applications with Python and the OpenAI API

 

Tanui Kipngetich Sila

September 19, 2026 • 20 min read

Retrieval-Augmented Generation can sound like a complicated artificial intelligence concept. Vector databases, embeddings, semantic search, retrieval pipelines, context windows, and language models all appear to be involved at once.

But there is a much easier way to understand it. Start with a simple Python program that sends some data to an AI model and asks questions about that data. Then gradually introduce retrieval.

That progression is important because a simple prompt-based application and a RAG application are not completely different ideas. RAG is essentially an extension of the same basic pattern.

In this tutorial, we will start with a small Python dataset, connect to the OpenAI API, discover which models our API key can access, ask questions about our data, identify the limitations of sending the entire dataset to the model, and then gradually build the architecture toward a real RAG system.

What We Are Building

By the end of the tutorial, the architecture we want will look like this:

                 YOUR DATA
                    |
                    v
               Documents
                    |
                    v
                 Chunks
                    |
                    v
               Embeddings
                    |
                    v
             Vector Storage
                    |
                    |
              USER QUESTION
                    |
                    v
             Query Embedding
                    |
                    v
           Similarity Search
                    |
                    v
            Relevant Context
                    |
                    v
          Question + Context
                    |
                    v
              OpenAI API
                    |
                    v
                Answer

Step 1: Create an OpenAI API Account

The first requirement is access to the OpenAI API. Your Python program communicates with the OpenAI platform through an API key.

You can access the official OpenAI developer platform here:

https://platform.openai.com/

From the platform, you can manage your API access and create an API key.

Open the OpenAI API Keys page

Security warning: Never publish your API key in GitHub repositories, screenshots, frontend JavaScript, blog posts, or public code. Treat the key like a password. If a key becomes exposed, rotate or revoke it.

Step 2: Install the Python SDK

Create a Python project and install the official OpenAI package.

pip install openai

You can verify that the package is installed:

pip show openai

Example output:

Name: openai
Version: 1.x.x
Summary: The official Python library for the OpenAI API

Step 3: Configure the API Key Safely

Avoid writing the API key directly into your Python file.

Instead, store it as an environment variable.

OPENAI_API_KEY=your_api_key_here

Then initialize the client:

from openai import OpenAI

client = OpenAI() 

Expected result:

Client initialized successfully.

Step 4: Find Out Which Models Your API Key Can Access

Before choosing a model, it is useful to ask the API which models are accessible to your account.

This prevents a common beginner problem: copying a model name from an example and discovering that the model is unavailable to the particular account, project, or API configuration.

The OpenAI model catalogue also provides information about model capabilities and supported tools.

from openai import OpenAI

client = OpenAI()

def list_available_models():

```
models = client.models.list()

print(
    f"Total models accessible: {len(models.data)}\n"
)

print("Accessible models:\n")

for model in sorted(
    models.data,
    key=lambda m: m.id
):
    print(model.id)
```

if **name** == "**main**":
list_available_models() 

Example output:

Total models accessible: 12

Accessible models:

gpt-5.6-luna
gpt-5.6-sol
... 

The exact list will depend on what is available to your API project. Do not treat the example output above as a guaranteed list.

Step 5: Test Your First API Request

Now that we know which models are available, select one that your account can access and make a simple request.

from openai import OpenAI

client = OpenAI()

model = "YOUR_AVAILABLE_MODEL"

response = client.responses.create(
model=model,
input="Explain what Python is in one sentence."
)

print(response.output_text) 

Example output:

Python is a high-level programming language known
for its readable syntax and broad range of applications.

Step 6: Give the Model Some Data

Now we move toward the example that makes the connection to RAG much clearer.

people = [
    {
        "name": "Jane Smith",
        "occupation": "Software Engineer",
        "description":
            "Works on backend systems and APIs."
    },
    {
        "name": "John Doe",
        "occupation": "Data Scientist",
        "description":
            "Works on machine learning projects."
    },
    {
        "name": "Mary Johnson",
        "occupation": "Product Manager",
        "description":
            "Leads software product development."
    }
]

The Python program now owns a small knowledge base. The model does not automatically have access to it. We have to explicitly provide the data.

Step 7: Convert the Data to JSON

JSON is a convenient format for putting structured Python data into a prompt.

import json

data = json.dumps(
people,
indent=2
)

print(data) 

Output:

[
  {
    "name": "Jane Smith",
    "occupation": "Software Engineer",
    "description": "Works on backend systems and APIs."
  },
  {
    "name": "John Doe",
    "occupation": "Data Scientist",
    "description": "Works on machine learning projects."
  }
]

Step 8: Ask Questions About the Data

Now we combine the data and the user's question into a single prompt.

def answer_question(question):

```
prompt = f"""
```

Here is a list of people:

{json.dumps(people, indent=2)}

Based on this data, answer the question:

{question}
"""

```
response = client.responses.create(
    model="YOUR_AVAILABLE_MODEL",
    input=prompt
)

return response.output_text
```

question = input("Ask your question: ")

answer = answer_question(question)

print(answer) 

Example input:

Who works in software?

Example output:

Jane Smith is a Software Engineer who works on
backend systems and APIs.

Mary Johnson is a Product Manager who leads
software product development. 

Step 9: Understand What We Have Built

At this point, we have a useful application, but we should be precise about what it is.

This is not yet a full RAG system.

We are performing what could be called prompt-based context injection. Every time the user asks a question, we put the entire dataset into the prompt.

That distinction is important because RAG introduces a retrieval stage between the user's question and the final prompt.

Step 10: Why the Simple Approach Eventually Breaks

Imagine expanding our employee list from three people to 100,000 people.

Every question would potentially send the entire dataset to the model.

100,000 records
       |
       v
Huge prompt
       |
       v
OpenAI API
       |
       v
Answer

This is where retrieval becomes useful. Instead of asking the model to inspect everything, we first find the information most likely to answer the question.

Step 11: Introduce Embeddings

An embedding is a numerical representation of text.

For example, these two questions use different words:

"How many credits do I need to graduate?"

"What is the required academic workload for graduation?"

A good embedding model represents the meaning of those sentences numerically so that semantic similarity can be measured.

This allows retrieval systems to search based on meaning rather than depending entirely on exact keyword matches.

Step 12: Calculate Similarity

Once the question and documents have been converted into vectors, we need to determine which vectors are closest.

One common measurement is cosine similarity.

import numpy as np

def cosine_similarity(a, b):

```
return np.dot(a, b) / (
    np.linalg.norm(a) *
    np.linalg.norm(b)
)
```

Example output:

Document 1: 0.82
Document 2: 0.41
Document 3: 0.27
Document 4: 0.12

Step 13: Retrieve the Top Results

Now we select the documents with the highest similarity scores.

top_indices = np.argsort(scores)[::-1][:3]

for index in top_indices:
print(documents[index]) 

Example output:

1. Graduation Requirements
2. MSc Programme Structure
3. Examination Regulations

Step 14: Build the Augmented Prompt

The retrieved documents now become the context for the language model.

context = "\n\n".join(retrieved_documents)

prompt = f"""
Answer the question using only the
provided context.

Context:
{context}

Question:
{question}

If the answer is not present in the
context, say that it is not available.
""" 

The resulting prompt conceptually looks like:

Context:
Students must complete 120 ECTS credits
before graduation.

Question:
How many credits do I need to graduate? 

Step 15: Generate the Final Answer

response = client.responses.create(
    model="YOUR_AVAILABLE_MODEL",
    input=prompt
)

print(response.output_text) 

Example output:

Students must complete 120 ECTS credits
before graduation.

Step 16: Add Sources to the Answer

A useful RAG system should ideally retain metadata about where every retrieved chunk came from.

{
    "text": "Students must complete 120 ECTS...",
    "source": "academic_regulations.pdf",
    "page": 42,
    "section": "Graduation Requirements"
}

Possible final response:

Students must complete 120 ECTS credits before graduation.

Source: Academic Regulations, page 42, Graduation Requirements

Different Real-World Scenarios

The same architecture can be used in many different applications.

1. University Assistant

Store course descriptions, academic regulations, examination rules, admission information, and student handbooks. Students can ask questions in natural language and receive answers grounded in university documents.

2. Company Knowledge Assistant

Index internal policies, onboarding documents, HR manuals, technical documentation, and standard operating procedures.

3. Customer Support Assistant

Retrieve troubleshooting instructions, product manuals, FAQs, and support articles before generating a response to a customer.

4. E-Commerce Assistant

Search product descriptions, specifications, inventory information, and policies before answering customer questions.

5. Research Assistant

Index research papers and technical documents so researchers can ask questions and retrieve relevant passages across a large collection of literature.

When Should You Use Simple Prompting?

Not every application needs RAG.

If you have ten records, a small JSON file, or a short configuration document, simply placing the data into the prompt may be the simplest solution.

Small data: Put the data directly into the prompt.

Large data: Retrieve the relevant information first.

RAG vs Fine-Tuning

RAG and fine-tuning solve different problems.

  • RAG: Retrieve external information and provide it to the model during a request.
  • Fine-tuning: Train a model further so that its behavior or task performance changes.

The Complete Mental Model

"RAG is a search layer placed in front of a language model."

User Question
      |
      v
Find relevant information
      |
      v
Retrieved Context
      |
      v
Question + Context
      |
      v
Language Model
      |
      v
Natural Language Answer

What to Build Next

The small employee example is intentionally simple. The next project should turn it into an actual document-question-answering system.

  1. Load PDF documents using Python.
  2. Extract their text.
  3. Split the text into chunks.
  4. Generate embeddings.
  5. Store embeddings in a vector database.
  6. Search the database when a user asks a question.
  7. Retrieve the most relevant chunks.
  8. Send those chunks to the OpenAI API.
  9. Return the answer together with document citations.

Conclusion

The easiest way to understand RAG is to build it progressively.

Start with a Python list. Serialize it. Put it into a prompt. Ask an AI model questions about it. Once the dataset becomes too large, introduce retrieval. Convert the documents and questions into embeddings, compare them, retrieve the relevant chunks, and provide those chunks to the model.

The language model is not replacing your database. It is not magically learning all your documents. Your application is finding the useful information and giving that information to the model at the right moment.

That is the fundamental idea behind Retrieval-Augmented Generation: retrieve the right information, augment the question with that information, and let the model generate the answer.

RAG Python OpenAI API Embeddings Vector Search AI Engineering
Tags:
#AI #OpenAI #Python #Retrieval-Augmented Generation #Machine Learning #API Integration

Discussion (0)

Leave a response ↓
No comments yet Be the first to share your thoughts on this article!

Leave a Comment

Your email address will not be published publicly. Name and email are optional (leave blank to comment anonymously).

Respectful & constructive discussion