A .NET-Powered RAG Console Application with Ollama

By · · AI Engineering

If you've been following the growth of Large Language Models and Retrieval-Augmented Generation applications, you've probably noticed a trend. The vast majority of tutorials, sample code, and applications are built in Python. Python's dominance in the AI/ML space is well-deserved given its rich ecosystem, but there's a solid alternative that gets overlooked in these conversations: .NET.

As someone who has worked extensively with both Python and .NET for AI applications, I'd like to share a recent project that shows why .NET deserves more attention for RAG systems.

A .NET-powered RAG console application

I recently built a complete RAG (Retrieval-Augmented Generation) system using .NET and Ollama, which lets you run LLM models locally. The architecture turned out clean, performance was solid, and the developer experience was smoother than I expected. Here's why I think .NET deserves a closer look for this kind of work.

5 reasons .NET works well for RAG and LLM applications

1. Strong type system and developer productivity

public interface IDocumentProcessor
{
    Task<List<DocumentChunk>> ProcessDocumentAsync(string filePath);
}

public class DocumentChunk
{
    public required string Text { get; set; }
    public required string DocumentName { get; set; }
    public int ChunkNumber { get; set; }
    public required float[] Embedding { get; set; }
}

One of the immediate benefits of .NET is its strong type system. The interface and class definitions above show how clean and self-documenting .NET code can be. As RAG systems grow more complex, compiler-enforced type checking catches errors early in development rather than at runtime.

2. Dependency injection built into the framework

var serviceProvider = new ServiceCollection()
    .AddLogging(configure => configure.AddConsole())
    .AddSingleton<IDocumentProcessor, PdfDocumentProcessor>()
    .AddSingleton<IVectorStore, SimpleVectorStore>()
    .AddSingleton<IEmbeddingService, OllamaEmbeddingService>()
    .AddSingleton<IChatService, OllamaChatService>()
    .AddSingleton<IRagService, RagService>()
    .BuildServiceProvider();

The dependency injection container in .NET makes it straightforward to build modular, testable applications. In the snippet above, we register all our services in the DI container, making them available throughout the application. This promotes loose coupling and makes it easy to swap implementations (for example, switching from Ollama to Azure OpenAI) without changing the consuming code.

3. Async/await model for handling I/O-bound operations

public async Task<string> GetAnswerAsync(string question)
{
    try
    {
        // Get embedding for the question
        var questionEmbedding = await _embeddingService.GetEmbeddingsAsync(question);

        // Retrieve similar documents
        var similarDocuments = await _vectorStore.GetSimilarDocumentsAsync(questionEmbedding, 3);

        // Concatenate the content of similar documents
        var context = new StringBuilder();
        foreach (var doc in similarDocuments)
        {
            context.AppendLine(
quot;From document: {doc.DocumentName}, chunk {doc.ChunkNumber}:"); context.AppendLine(doc.Text); context.AppendLine(); } // Get answer from chat service string answer = await _chatService.GetResponseAsync(question, context.ToString()); return answer; } catch (Exception ex) { return
quot;Error: {ex.Message}"; } }

The async/await pattern in .NET handles I/O-bound operations cleanly, which matters for RAG systems that make API calls to the LLM service, run database queries, and access the file system. The code stays readable while using system resources efficiently.

4. Performance and resource efficiency

While Python has the simpler syntax and the richer ML ecosystem, .NET has real performance advantages:

For production RAG systems that need to handle many concurrent requests or run within resource constraints, these matter. In my testing, the .NET implementation handled PDF processing and embedding generation quickly with minimal resource usage.

5. Enterprise-ready ecosystem

Many organizations already have significant investments in .NET infrastructure. Building RAG systems with .NET allows for:

The implementation details

My .NET RAG implementation follows a clean architecture approach with clearly defined interfaces:

The entire system is wired together with dependency injection, so each component is testable and replaceable.

A note on ecosystem maturity

Python's ecosystem for machine learning and AI is more mature, no question. Libraries like Hugging Face Transformers, LangChain, and LlamaIndex are industry standards. But the .NET ecosystem is catching up:

When to choose .NET for your RAG system

I'm not suggesting .NET should replace Python in the AI space. But there are scenarios where it makes good sense:

Step-by-step guide to building your own .NET RAG system

Here's a walkthrough for building your own .NET-powered RAG application:

1. Project setup

# Create a new console application
dotnet new console -n dotnet_console_rag_ollama
cd dotnet_console_rag_ollama

# Add necessary packages
dotnet add package Microsoft.Extensions.DependencyInjection
dotnet add package Microsoft.Extensions.Logging
dotnet add package Microsoft.Extensions.Logging.Console
dotnet add package itext7 # For PDF processing
dotnet add package Newtonsoft.Json

2. Define your interfaces

Start by creating clear interfaces that define the responsibilities of each component:

public interface IDocumentProcessor
{
    Task<List<DocumentChunk>> ProcessDocumentAsync(string filePath);
}

public interface IEmbeddingService
{
    Task<float[]> GetEmbeddingsAsync(string text);
}

public interface IVectorStore
{
    Task AddDocumentAsync(DocumentChunk document);
    Task<List<DocumentChunk>> GetSimilarDocumentsAsync(float[] queryEmbedding, int topK);
}

public interface IChatService
{
    Task<string> GetResponseAsync(string question, string context);
}

public interface IRagService
{
    Task ProcessDocumentsAsync(string folderPath);
    Task<string> GetAnswerAsync(string question);
}

3. Implement document processing

Create a PDF document processor that extracts text from PDFs and chunks it:

public class PdfDocumentProcessor : IDocumentProcessor
{
    private readonly ILogger<PdfDocumentProcessor> _logger;
    private readonly IEmbeddingService _embeddingService;
    private const int MaxChunkSize = 1000; // characters per chunk

    public PdfDocumentProcessor(ILogger<PdfDocumentProcessor> logger, IEmbeddingService embeddingService)
    {
        _logger = logger;
        _embeddingService = embeddingService;
    }

    public async Task<List<DocumentChunk>> ProcessDocumentAsync(string filePath)
    {
        _logger.LogInformation(
quot;Processing PDF: {filePath}"); // Extract text from PDF using iText7 string extractedText = ExtractTextFromPdf(filePath); // Split text into chunks var textChunks = ChunkText(extractedText, MaxChunkSize); List<DocumentChunk> documentChunks = new List<DocumentChunk>(); // Create document chunks with embeddings for (int i = 0; i < textChunks.Count; i++) { var embedding = await _embeddingService.GetEmbeddingsAsync(textChunks[i]); documentChunks.Add(new DocumentChunk { Text = textChunks[i], DocumentName = Path.GetFileName(filePath), ChunkNumber = i + 1, Embedding = embedding }); } return documentChunks; } // Implementation details for ExtractTextFromPdf and ChunkText methods... }

4. Implement embedding service (Ollama integration)

Create a service that connects to Ollama to generate embeddings:

public class OllamaEmbeddingService : IEmbeddingService
{
    private readonly ILogger<OllamaEmbeddingService> _logger;
    private readonly HttpClient _httpClient;
    private const string EmbeddingModel = "nomic-embed-text";
    private const string OllamaBaseUrl = "http://localhost:11434/api";

    public OllamaEmbeddingService(ILogger<OllamaEmbeddingService> logger)
    {
        _logger = logger;
        _httpClient = new HttpClient();
    }

    public async Task<float[]> GetEmbeddingsAsync(string text)
    {
        try
        {
            var request = new
            {
                model = EmbeddingModel,
                prompt = text
            };

            var response = await _httpClient.PostAsJsonAsync(
                
quot;{OllamaBaseUrl}/embeddings", request); // Process response to extract embeddings // ... return embeddings; } catch (Exception ex) { _logger.LogError(
quot;Error getting embeddings: {ex.Message}"); throw; } } }

5. Implement vector store

Create a simple in-memory vector store with cosine similarity search:

public class SimpleVectorStore : IVectorStore
{
    private readonly List<DocumentChunk> _documents = new List<DocumentChunk>();
    private readonly ILogger<SimpleVectorStore> _logger;

    public SimpleVectorStore(ILogger<SimpleVectorStore> logger)
    {
        _logger = logger;
    }

    public Task AddDocumentAsync(DocumentChunk document)
    {
        _documents.Add(document);
        return Task.CompletedTask;
    }

    public Task<List<DocumentChunk>> GetSimilarDocumentsAsync(float[] queryEmbedding, int topK)
    {
        // Calculate cosine similarity for each document
        var similarities = _documents
            .Select(doc => new
            {
                Document = doc,
                Similarity = CosineSimilarity(queryEmbedding, doc.Embedding)
            })
            .OrderByDescending(x => x.Similarity)
            .Take(topK)
            .Select(x => x.Document)
            .ToList();

        return Task.FromResult(similarities);
    }

    private float CosineSimilarity(float[] vector1, float[] vector2)
    {
        // Implementation of cosine similarity calculation
        // ...
    }
}

6. Implement chat service

Create a service that sends prompts to Ollama's LLM:

public class OllamaChatService : IChatService
{
    private readonly ILogger<OllamaChatService> _logger;
    private readonly HttpClient _httpClient;
    private const string ChatModel = "llama3:8b";
    private const string OllamaBaseUrl = "http://localhost:11434/api";

    public OllamaChatService(ILogger<OllamaChatService> logger)
    {
        _logger = logger;
        _httpClient = new HttpClient();
    }

    public async Task<string> GetResponseAsync(string question, string context)
    {
        try
        {
            string systemPrompt = "You are a helpful assistant that answers questions based on the provided context.";
            string prompt = 
quot;Context:\n{context}\n\nQuestion: {question}\n\nAnswer:"; var request = new { model = ChatModel, prompt = prompt, system = systemPrompt, stream = false }; var response = await _httpClient.PostAsJsonAsync(
quot;{OllamaBaseUrl}/generate", request); // Process response to extract answer // ... return answer; } catch (Exception ex) { _logger.LogError(
quot;Error getting response: {ex.Message}"); return
quot;Error: {ex.Message}"; } } }

7. Implement the RAG service

Create the main service that orchestrates the entire RAG process:

public class RagService : IRagService
{
    private readonly ILogger<RagService> _logger;
    private readonly IDocumentProcessor _documentProcessor;
    private readonly IVectorStore _vectorStore;
    private readonly IEmbeddingService _embeddingService;
    private readonly IChatService _chatService;

    public RagService(
        ILogger<RagService> logger,
        IDocumentProcessor documentProcessor,
        IVectorStore vectorStore,
        IEmbeddingService embeddingService,
        IChatService chatService)
    {
        _logger = logger;
        _documentProcessor = documentProcessor;
        _vectorStore = vectorStore;
        _embeddingService = embeddingService;
        _chatService = chatService;
    }

    public async Task ProcessDocumentsAsync(string folderPath)
    {
        // Process all PDF files in the specified folder
        // ...
    }

    public async Task<string> GetAnswerAsync(string question)
    {
        // Get embedding for the question
        var questionEmbedding = await _embeddingService.GetEmbeddingsAsync(question);

        // Retrieve similar documents
        var similarDocuments = await _vectorStore.GetSimilarDocumentsAsync(questionEmbedding, 3);

        // Prepare context from similar documents
        // ...

        // Get answer from LLM
        string answer = await _chatService.GetResponseAsync(question, context);

        return answer;
    }
}

8. Wire everything together

In your Program.cs, set up the dependency injection container and create the main application loop:

static async Task Main(string[] args)
{
    // Set up dependency injection
    var serviceProvider = new ServiceCollection()
        .AddLogging(configure => configure.AddConsole())
        .AddSingleton<IDocumentProcessor, PdfDocumentProcessor>()
        .AddSingleton<IVectorStore, SimpleVectorStore>()
        .AddSingleton<IEmbeddingService, OllamaEmbeddingService>()
        .AddSingleton<IChatService, OllamaChatService>()
        .AddSingleton<IRagService, RagService>()
        .BuildServiceProvider();

    var ragService = serviceProvider.GetRequiredService<IRagService>();

    // Process documents
    await ragService.ProcessDocumentsAsync("./Documents");

    // Interactive question loop
    while (true)
    {
        Console.Write("\nYour question (type 'exit' to quit): ");
        string question = Console.ReadLine() ?? string.Empty;

        if (question.ToLower() == "exit") break;

        string answer = await ragService.GetAnswerAsync(question);
        Console.WriteLine(
quot;\nAnswer: {answer}"); } }

9. Set up and configure Ollama

Before running your application, you need to set up Ollama:

  1. Install Ollama from https://ollama.ai

    • For macOS: Download and install the .dmg file
    • For Windows: Download and run the installer
    • For Linux: Use the install script curl -fsSL https://ollama.com/install.sh | sh
  2. Start the Ollama service:

    ollama serve
    

    This will start the Ollama API server on http://localhost:11434

  3. Pull the required models (in a new terminal window):

    # Pull the embedding model
    ollama pull nomic-embed-text
    
    # Pull the chat model
    ollama pull llama3:8b
    
  4. Verify the models are correctly installed:

    # List all available models
    ollama list
    

    You should see both "nomic-embed-text" and "llama3:8b" in the list.

  5. Test the embedding model:

    # Test embedding generation
    curl -X POST http://localhost:11434/api/embeddings -d '{
      "model": "nomic-embed-text",
      "prompt": "This is a test."
    }' | head
    

    You should see a JSON response with an "embedding" array containing vector values.

  6. Test the chat model:

    # Test text generation
    curl -X POST http://localhost:11434/api/generate -d '{
      "model": "llama3:8b",
      "prompt": "What is retrieval-augmented generation?",
      "stream": false
    }' | jq '.response'
    

    You should receive a helpful response explaining RAG. If you don't have jq installed, you can omit the | jq '.response' part.

Additional Ollama commands that might be useful:

# Check Ollama version
ollama --version

# Remove a model
ollama rm model-name

# Show model information
ollama show nomic-embed-text

# Update a model to the latest version
ollama pull llama3:8b:latest

10. Test your application

Build and run your application:

# Build the application
dotnet build

# Create a Documents directory if it doesn't exist
mkdir -p Documents

# Add some PDF documents to test with
# cp /path/to/your/documents/*.pdf ./Documents/

# Run the application
dotnet run

Place some PDF documents in the ./Documents folder and start asking questions related to the content!

Troubleshooting common issues

Performance optimization

For better performance:

Conclusion

The next time you plan a RAG system or other LLM-powered application, it's worth considering .NET before defaulting to Python. My implementation shows that .NET is a solid foundation for AI applications with clean, maintainable code.

The complete source code for this .NET RAG console application is available on my GitHub. repo: https://github.com/encryptedtouhid/dotnet_console_rag_ollama

Feel free to check it out, contribute, or adapt it for your own projects. Happy Coding!