If you have been developing Rails applications for years, there’s a good chance you’ve used:
bin/rails credentials:edit
hundreds of times.
You probably know that Rails stores encrypted credentials in:
config/credentials.yml.enc
and keeps the encryption key separately in:
config/master.key
But did you know that Rails can make:
git diff
show the decrypted, human-readable changes to credentials.yml.enc?
I recently discovered this while working on a Rails 8.1.3.1 application and it was one of those:
“I’ve been using Rails every day for years, and I didn’t know Rails could do this!”
moments.
Let’s see how it works.
First: What is credentials.yml.enc?
Rails encrypted credentials allow us to keep secrets such as:
openai:
api_key: ...
or:
aws:
access_key_id: ...
secret_access_key: ...
inside:
config/credentials.yml.enc
The file is encrypted.
The encryption key is stored separately in:
config/master.key
Rails documentation explicitly states that the encrypted credentials file can be stored in version control as long as the master key remains secure. (Ruby on Rails Guides)
So our repository can contain:
config/
├── credentials.yml.enc ← encrypted, safe to commit
└── master.key ← secret, NEVER commit
Editing Rails Credentials
Normally we edit credentials with:
bin/rails credentials:edit
Rails decrypts the credentials, opens them in your configured editor, and encrypts them again when you save.
Rails then ensures the Git diff driver is configured to use:
bin/rails credentials:diff
Rails’ application generator includes this credentials diff enrollment as part of application setup and Rails 7.0 already contained the credentials diffing implementation. (Gem)
So this isn’t actually an 8.1-only feature.
That’s an important distinction.
Is This New in Rails 8.1?
No – and this is an important correction.
The encrypted credentials diff functionality existed before Rails 8.1.
For example, Rails 7.0 already had the credentials:diff implementation, and Rails 7.2’s application generator also enrolled projects in credentials diffing. (Gem)
Rails has supported decrypted Git diffs for encrypted credentials for several versions and Rails 8.x continues to build on the credentials tooling.
Rails 8.1 does introduce other useful credentials functionality. For example, Rails 8.1 added command-line credential fetching, which can be useful for deployment tooling such as Kamal. (Ruby on Rails Guides)
The master key should not be committed. Rails’ security guide explicitly recommends keeping the master key safe and out of version control. (Ruby on Rails Guides)
One Thing to Remember
The decrypted content can appear in your local terminal output.
The difference is a great way to understand what’s really happening.
Quick Reference
# Edit credentials
bin/rails credentials:edit
# Enroll project in credential diffing
bin/rails credentials:diff --enroll
# Normal readable diff
git diff
# Show the actual encrypted file diff
git diff --no-textconv -- config/credentials.yml.enc
# Inspect Git's configuration
git config --show-origin --get-regexp 'diff|textconv|filter'
# Check Git attributes
git check-attr diff -- config/credentials.yml.enc
Security rule:
Y config/credentials.yml.enc → commit it
X config/master.key → NEVER commit it
Rails’ official security guide confirms that encrypted credentials can be stored in version control while the master key must remain protected. (Ruby on Rails Guides)
Ruby 4.0 introduced two fascinating runtime capabilities:
Ractors, significantly improved for parallel execution
Ruby Box, an experimental mechanism for isolating definitions inside one Ruby process
For a Rails developer, the obvious question is:
Can I take my existing Rails application and simply add Ractors and Ruby Box to make it faster or more scalable?
The answer is not yet that simple.
Ractors can be extremely useful for carefully isolated CPU-heavy work, but a conventional Rails application is deeply interconnected through global state, constants, classes, ActiveSupport, ActiveRecord, gems, configuration and caches.
Ruby Box is a completely different concept. It is not primarily a parallelism mechanism. It provides in-process isolation of definitions and loaded code, with potential applications such as running multiple application versions in one Ruby process. Ruby 4.0 documents it as experimental.
Let’s look at both from a Rails perspective.
1. First: what problem does a Ractor solve?
A normal Ruby thread looks roughly like this:
Rails process
│
├── Thread 1
├── Thread 2
├── Thread 3
└── Thread 4
│
└── same Ractor / same GVL
Threads within a Ractor still share that Ractor’s GVL, so they don’t execute Ruby code in parallel with one another.
Ractors change the model:
Rails process
│
├── Ractor A ── GVL ── Thread(s)
│
├── Ractor B ── GVL ── Thread(s)
│
└── Ractor C ── GVL ── Thread(s)
Different Ractors can execute Ruby code in parallel on different CPU cores. Ruby 4.0 also reduced internal contention and introduced Ractor::Port for communication.
That makes Ractors especially interesting for CPU-bound work.
2. What should NOT be your first Ractor experiment?
Suppose you have:
class ReportsController < ApplicationController
def show
@report = Report.generate
end
end
It is tempting to write:
def show
r = Ractor.new do
Report.generate
end
@report = r.value
end
This is exactly the kind of approach that exposes the biggest problem.
A Rails application has a huge amount of shared framework state.
For example:
Rails
│
├── ActiveSupport
├── ActiveRecord
├── Zeitwerk
├── configuration
├── caches
├── logging
├── autoloading
├── class/module definitions
└── gems
Ractors deliberately restrict access to non-shareable objects across Ractors.
The Ruby documentation says that most objects are unshareable and communication between Ractors is intended to happen through shareable objects or message passing.
That makes a normal Rails application a poor candidate for simply wrapping arbitrary Rails calls inside Ractor.new.
There has also been a real Rails issue demonstrating Ractor::IsolationError when attempting to instantiate or use Rails application state from a non-main Ractor.
3. The better idea: use Ractors around isolated computation
Instead of:
Ractor
↓
Entire Rails application
think:
Rails
│
├── request
│
├── database work
│
└── isolated CPU calculation
↓
Ractor
For example, imagine a report containing millions of values.
class ReportCalculator
def self.calculate(numbers)
numbers.sum { |n| expensive_calculation(n) }
end
def self.expensive_calculation(n)
# CPU-heavy calculation
n ** 3
end
end
You could partition the data:
chunks = numbers.each_slice(10_000).to_a
ractors = chunks.map do |chunk|
Ractor.new(chunk) do |values|
values.sum { |n| n ** 3 }
end
end
result = ractors.sum(&:value)
The important architectural boundary is:
Rails
│
│ plain data
▼
Ractor 1 ── CPU work ──┐
Ractor 2 ── CPU work ──┼──→ results
Ractor 3 ── CPU work ──┘
│
▼
Rails
This is much more promising.
The Ractors don’t need to manipulate:
ActiveRecord::Relation
Rails.application
ActiveSupport::Cache
Controller
request
response
They receive isolated data and return isolated results.
4. A practical Rails use case: analytics
Imagine:
orders=Order
.where(created_at:30.days.ago..)
.pluck(:amount)
The database query happens normally.
Then:
chunks = orders.each_slice(50_000).to_a
ractors = chunks.map do |chunk|
Ractor.new(chunk) do
{
total: chunk.sum,
average: chunk.sum.to_f / chunk.length
}
end
end
results = ractors.map(&:value)
total = results.sum { |r| r[:total] }
The database remains Rails’ responsibility.
The CPU-heavy aggregation becomes parallel work.
That is the mental model I’d recommend:
Use Rails for orchestration; use Ractors for isolated computation.
5. Another good candidate: document/image processing
Suppose your application performs CPU-heavy transformations:
PDF
↓
parse
↓
transform
↓
calculate
↓
generate result
Instead of letting one Ruby execution stream process everything:
Rails
│
└── CPU-heavy processing
you can potentially build:
Rails
│
Job / Service
│
┌───────────┼───────────┐
▼ ▼ ▼
Ractor Ractor Ractor
│ │ │
file A file B file C
└──────────┬──────────┘
▼
result
The same principle applies to:
compression
large JSON transformations
encryption-related computation
parsing
ranking/scoring
simulations
large in-memory calculations
The exact benefit depends heavily on whether the work is CPU-bound and whether the cost of copying/moving data outweighs the parallelism benefit.
Ruby’s Ractor documentation explicitly notes that unshareable objects may be copied or moved between Ractors, so data-transfer overhead must be considered.
6. Ractor is not a replacement for ActiveJob
It is important not to confuse these abstractions.
For example:
SomeJob.perform_later(order.id)
and:
Ractor.new(...)
solve different problems.
ActiveJob/Sidekiq/GoodJob/etc. solve background job execution and process-level application architecture.
Ractors solve parallel execution inside one Ruby process.
You could potentially combine them:
Rails
│
▼
Background Job
│
▼
Ruby Process
│
├── Ractor 1
├── Ractor 2
├── Ractor 3
└── Ractor 4
But this is an advanced optimization, not the default architecture.
7. So where does Ruby Box fit?
Ruby Box is much less about CPU parallelism.
Its purpose is definition isolation.
Suppose you have:
classUser
defrole
"admin"
end
end
Now imagine loading another piece of code that reopens User:
classUser
defrole
"guest"
end
end
Normally, that’s a global change to the Ruby process.
Ruby Box lets those definitions exist in separate boxes.
Conceptually:
Ruby process
│
├── Main box
│ └── User#role → "admin"
│
└── Box B
└── User#role → "guest"
Ruby’s documentation describes this as isolation of class/module definitions, monkey patches, constants, global/class variables and loaded Ruby/native libraries.
This is a very different problem from Ractors.
8. A simple Ruby Box example
Ruby Box must be enabled at process startup:
RUBY_BOX=1 ruby app.rb
Setting the variable after Ruby has already started does not enable it.
Then:
box=Ruby::Box.new
box.require("./legacy_user.rb")
Suppose legacy_user.rb contains:
classUser
defrole
"legacy"
end
end
The definition is loaded into the box.
Conceptually:
Main box
│
└── User
Legacy box
│
└── User
└── role → "legacy"
The definition in the box is isolated from the corresponding definition in other boxes. Ruby’s documentation demonstrates this with constants, classes and methods.
9. This has a fascinating Rails use case: blue-green application versions
This is one of the use cases Ruby itself proposes.
Imagine:
One Ruby process
┌─────────────────────────────┐
│ Ruby Process │
│ │
│ Box A │
│ Rails App v1 │
│ │
│ Box B │
│ Rails App v2 │
└─────────────────────────────┘
Ruby 4.0 explicitly lists running web-app boxes in parallel as a potential blue-green deployment use case.
Theoretically, this gives you the ability to have:
/app-v1
/app-v2
loaded in separate definition environments inside the same Ruby process.
Then requests could be directed to:
traffic
│
├──→ Box A
│
└──→ Box B
This could eventually enable interesting deployment and migration strategies.
But there is a huge caveat.
10. Ruby Box is experimental
This is not currently something I’d take into a normal Rails production deployment simply because Ruby 4 has it.
The official documentation lists known issues, including:
Keep Rails state out of the Ractors wherever possible.
Design explicit boundaries:
input= {
values:values,
options:options
}
rather than:
ractor=Ractor.newdo
Order.where(...)
end
The first is an isolated computation.
The second makes the Ractor responsible for Rails state.
That’s where the complexity explodes.
13. How I would introduce Ractors into an existing Rails app
Start with one measurable CPU bottleneck.
For example:
Before
request
↓
large calculation
↓
1 CPU core
↓
response
Then extract:
classPricingCalculator
defself.calculate(input)
# pure Ruby calculation
end
end
Make it as pure as possible:
result=PricingCalculator.calculate(
prices:prices,
rules:rules
)
Then experiment with:
Ractor.new(input) do |data|
PricingCalculator.calculate(data)
end
Benchmark both:
single-threaded
vs
multiple Ractors
Don’t assume parallelism automatically means faster execution.
You need to measure:
CPU time
wall-clock time
memory usage
object copying
Ractor startup
throughput
latency
14. Rails architecture: where each feature fits
A useful mental model is:
Rails
│
┌────────────────┼─────────────────┐
│ │ │
▼ ▼ ▼
Web Data Jobs
│ │ │
└────────────────┴────────┐ │
▼ ▼
Application services
│
CPU-heavy workload
│
┌─────┴─────┐
▼ ▼
Ractor Ractor
Ruby Box sits at a different architectural layer:
Ruby Process
│
├── Main / Application Box
│
├── Application Box A
│
└── Application Box B
So:
Ractor = parallel execution
while:
Ruby Box = definition/environment isolation
They solve different problems.
15. The big Rails limitation today
This is the part worth remembering.
A conventional Rails application is built around a substantial amount of shared application state.
That doesn’t fit naturally with Ractor’s isolation model.
There has been an explicit Rails issue requesting Ractor support and that issue was closed as “not planned.” The discussion showed Ractor::IsolationError arising from Rails class-level state.
That architecture gives you a much better chance of benefiting from parallel Ruby.
For Ruby Box:
Development/testing first
↓
isolated definitions
↓
experimentation
↓
specialized deployment scenarios
rather than immediately attempting:
"Let's run the whole Rails app in 10 Ruby Boxes."
Final takeaway
Ruby 4 did something more interesting than simply making threads faster.
It is giving Ruby developers more explicit runtime tools:
Ractor
↓
parallel Ruby computation
Ruby Box
↓
isolated Ruby definitions
YJIT / ZJIT
↓
faster execution
GC/runtime improvements
↓
less overhead
For Rails, however, the winning strategy isn’t:
“Convert Rails to Ractors.”
It is:
“Keep Rails responsible for application orchestration and isolate carefully chosen CPU-heavy computations behind Ractor boundaries.”
And Ruby Box is even more experimental.
Its long-term Rails potential may actually be more architectural than performance-oriented: isolating application versions, tests, plugins, or dependency environments inside one Ruby process.
The exciting thing is that Ruby 4 gives us primitives that make these designs possible.
The engineering challenge is deciding where the boundary belongs.
That is ultimately the same lesson we’ve been following through this series:
Ruby gives us abstractions. Understanding the runtime lets us decide when to cross them.
A useful next post in this series would be “Ractor vs Thread vs Process in Rails: when should a senior Rails developer choose each?” – with real benchmarks, memory/CPU trade-offs and a Sidekiq/Puma/Ractor architecture comparison.
Database transactions are one of those Rails features that appear simple:
User.transaction do
# database operations
end
But at a senior Rails engineering level, transactions are much more than wrapping a few save! calls in a block.
They affect data consistency, concurrency, failure handling, callbacks, database locking, nested service objects, multiple databases and even how external systems should be triggered.
A strong Rails developer should understand not only how to start a transaction, but also:
which transaction API to use
how class-level and instance-level transactions differ
what actually gets rolled back
how nested transactions work
when requires_new is necessary
how savepoints work
how save and destroy already use transactions internally
how after_commit differs from after_save
how to handle database exceptions safely
how transactions interact with locks and isolation levels
what Rails transactions cannot protect
how to design transaction boundaries in service objects
This article approaches transactions from that perspective.
What Is a Database Transaction?
A transaction groups multiple database operations into a single unit of work.
Conceptually:
BEGIN
operation 1
operation 2
operation 3
COMMIT
If something fails:
BEGIN
operation 1
operation 2
operation 3 <-- failure
ROLLBACK
The important property is atomicity:
Either all database changes become permanent, or none of them do.
Account.transaction do
sender.withdraw!(100)
receiver.deposit!(100)
end
We don’t want this situation:
sender: -100
receiver: 0
If the deposit fails, the withdrawal must also disappear.
The Basic Rails Transaction Block
The most common syntax is:
ActiveRecord::Base.transaction do
user.save!
profile.save!
audit_log.save!
end
If any operation raises an exception, Rails rolls back the transaction.
In a modern Rails application, however, I generally prefer:
ApplicationRecord.transaction do
user.save!
profile.save!
audit_log.save!
end
Why?
ApplicationRecord represents the application’s Active Record hierarchy and is generally a clearer boundary than directly referencing ActiveRecord::Base.
Why Do We Need Transactions?
Imagine an order creation workflow:
order = Order.create!
payment = Payment.create!
order.update!(status: "paid")
InventoryItem.create!(...)
Without a transaction, failure halfway through could leave:
Order ✅
Payment ✅
Order paid ✅
Inventory ❌
The application is now inconsistent.
Instead:
ApplicationRecord.transaction do
order = Order.create!
payment = Payment.create!
order.update!(status: "paid")
InventoryItem.create!
end
order.transaction do
order.update!(status: "processing")
order.create_payment!
end
This reads naturally as:
Perform these operations as one transaction around this order.
However, I would usually use the class-level transaction in service objects, because the service boundary is about a unit of work rather than a specific model.
Service Object Transaction Boundary
This is usually my preferred architecture for complex workflows.
class CheckoutService
def call(user:, cart:)
Order.transaction do
order = create_order(user, cart)
charge_payment(order)
reserve_inventory(order)
order.update!(status: "confirmed")
order
end
end
private
def create_order(user, cart)
Order.create!(
user: user,
total: cart.total
)
end
def charge_payment(order)
Payment.create!(
order: order,
amount: order.total
)
end
def reserve_inventory(order)
# ...
end
end
This gives the workflow one explicit transaction boundary.
That is much easier to reason about than having every individual method start its own transaction.
is already transaction-protected at the database-operation level.
But that doesn’t mean you don’t need explicit transactions.
Consider:
user.save!
profile.save!
Each operation is protected individually.
That does not give you:
user + profile = atomic unit
You need:
User.transaction do
user.save!
profile.save!
end
So the distinction is:
save!
↓
protect this persistence operation
transaction do
↓
protect this entire workflow
Transaction and Exceptions
The normal rollback mechanism is an exception.
User.transaction do
user.save!
profile.save!
raise "Something failed"
end
The transaction rolls back.
Rails then propagates the exception to the caller.
That means this pattern is common:
begin
User.transaction do
create_user!
create_profile!
end
rescue StandardError => e
Rails.logger.error(e.message)
raise
end
The important design principle is:
Don’t use transactions as a replacement for error handling.
A transaction determines what happens to database state.
Your application code still needs to determine what happens to the failure.
ActiveRecord::Rollback
Rails provides a special exception:
ActiveRecord::Rollback
Example:
User.transaction do
user.update!(status: "processing")
unless payment_valid?
raise ActiveRecord::Rollback
end
user.update!(status: "confirmed")
end
The transaction rolls back, but ActiveRecord::Rollback is specifically handled by Rails and isn’t propagated like a normal exception. (Ruby on Rails API)
This makes it useful when you intentionally want:
rollback database changes
+
don't treat this as an application exception
For example:
Order.transaction do
order.update!(status: "processing")
raise ActiveRecord::Rollback unless inventory_available?
end
raise vs ActiveRecord::Rollback
Compare:
raisePaymentError
with:
raiseActiveRecord::Rollback
The semantic difference is significant.
Normal exception
raisePaymentError
Result:
ROLLBACK
↓
exception propagates
↓
caller can rescue it
ActiveRecord::Rollback
raiseActiveRecord::Rollback
Result:
ROLLBACK
↓
rollback is internally handled
↓
execution exits transaction
Therefore, don’t blindly replace business exceptions with ActiveRecord::Rollback.
Nested Transactions
Consider:
ApplicationRecord.transaction do
user.save!
ApplicationRecord.transaction do
profile.save!
end
end
You might assume there are two independent transactions:
BEGIN
user
BEGIN
profile
COMMIT
COMMIT
That’s not generally how Rails works.
By default, a nested transaction joins the existing transaction. Most databases don’t provide true nested transactions, so Rails emulates subtransactions with savepoints where necessary. (Ruby on Rails API)
Conceptually:
BEGIN
user
profile
COMMIT
The Nested Rollback Surprise
Consider this:
User.transactiondo
User.create!(name:"A")
User.transactiondo
User.create!(name:"B")
raiseActiveRecord::Rollback
end
end
Many developers expect:
A -> committed
B -> rolled back
But because the nested transaction joins the parent transaction, the rollback exception is handled by the inner transaction boundary and the outer transaction can still commit.
The result can be:
A -> committed
B -> committed
This behavior is documented by Rails and is one of the most important transaction gotchas to understand. (Ruby on Rails API)
requires_new: true
When you need an actual nested transactional boundary, use:
requires_new:true
Example:
ApplicationRecord.transactiondo
user=User.create!
ApplicationRecord.transaction(requires_new:true) do
AuditLog.create!
raiseActiveRecord::Rollback
end
end
Now the inner transaction gets its own savepoint.
Conceptually:
BEGIN
User
SAVEPOINT
AuditLog
ROLLBACK TO SAVEPOINT
COMMIT
Result:
User ✅
AuditLog ❌
Rails uses database save points to emulate nested transactions on databases that do not support true nested transactions.
When Should You Use requires_new?
It is especially useful when you have an inner operation that should be isolated from the outer workflow.
For example:
Order.transactiondo
create_order!
Order.transaction(requires_new:true) do
create_optional_audit_record!
end
finalize_order!
end
The inner operation can fail without necessarily destroying the outer work.
This is particularly useful in reusable service objects.
Suppose:
classAuditService
defself.record!(event)
AuditLog.transaction(requires_new:true) do
AuditLog.create!(event:event)
end
end
end
That service can be called either:
AuditService.record!("user_created")
or inside another transaction:
User.transactiondo
user.save!
AuditService.record!("user_created")
end
The requires_new boundary gives the service an explicit savepoint when called inside an existing transaction.
Transaction Isolation Levels
Transactions don’t only provide atomicity.
They also influence how concurrent transactions can see and modify data.
Rails allows:
Account.transaction(isolation::serializable) do
# critical operation
end
Supported isolation levels include:
:read_uncommitted
:read_committed
:repeatable_read
:serializable
Support and semantics depend on the database adapter. Rails documents these options and notes that isolation cannot generally be changed while joining an existing transaction or creating a nested savepoint transaction.
For example:
Order.transaction(isolation::serializable) do
order=Order.find(order_id)
order.update!(
status:"confirmed"
)
end
This can be appropriate for highly concurrent business operations, but it isn’t something I would enable casually.
Higher isolation can increase contention and retry requirements.
A senior engineer should ask:
What consistency guarantee does this business operation actually require?
rather than:
Which isolation level sounds safest?
Transactions + Row Locking
Transactions become particularly powerful when combined with pessimistic locking.
Suppose two requests attempt to update the same account simultaneously.
account.with_lockdo
account.update!(
balance:account.balance-100
)
end
with_lock is a useful Rails shortcut: it starts a transaction, reloads the record using a row lock, and then executes the block. Rails also allows transaction options such as requires_new, isolation, and joinable to be passed to with_lock. (Ruby on Rails API)
The important point is that transaction + locking solves a different class of problem from transaction alone.
A transaction provides atomicity.
A lock helps control concurrent access.
with_lock vs transaction
Compare:
account.transactiondo
account.update!(balance:...)
end
with:
account.with_lockdo
account.update!(balance:...)
end
The first provides transactional atomicity.
The second provides:
transaction
+
record reload
+
row-level locking
Use with_lock when concurrency around a particular row is part of the problem.
after_save vs after_commit
This distinction becomes extremely important when transactions are involved.
Consider:
classOrder<ApplicationRecord
after_save:publish_order
defpublish_order
EventBus.publish(id)
end
end
Suppose:
Order.transactiondo
order.update!(status:"paid")
raise"Something failed"
end
The after_save callback can run while the transaction is still in progress.
The database transaction can subsequently roll back.
Now your external system might have received:
Order paid
while your database says:
Order unpaid
That’s dangerous.
Use after_commit for External Side Effects
Rails provides:
classOrder<ApplicationRecord
after_commit:publish_order
private
defpublish_order
EventBus.publish(id)
end
end
Now the external operation happens only after the database transaction has successfully committed.
Rails explicitly recommends transaction callbacks such as after_commit when interacting with systems outside the database transaction. (Ruby on Rails Guides)
For narrower cases:
after_create_commit:publish_order
after_update_commit:publish_order
after_destroy_commit:remove_from_search
These are convenient aliases provided by Rails.
Per-Transaction Callbacks
Modern Rails also allows callbacks to be registered directly against a transaction.
For example:
Order.transactiondo |transaction|
order.update!(status:"confirmed")
transaction.after_commitdo
NotificationService.notify_order_confirmed(order)
end
end
This is interesting because the callback is associated with the unit of work, rather than with the model lifecycle.
Rails supports transaction-level callbacks such as:
transaction.before_commit
transaction.after_commit
transaction.after_rollback
This can be cleaner for domain/service-oriented workflows where you don’t want the model itself to know about notification behavior.
ActiveRecord.after_all_transactions_commit
Another useful modern Rails API is:
ActiveRecord.after_all_transactions_commitdo
NotificationService.notify(...)
end
This is useful when code may be invoked from either inside or outside a transaction.
Rails guarantees that the callback runs after all currently open transactions have successfully committed. If any transaction rolls back, the callback isn’t executed.
This can be particularly useful in reusable application services.
Article.current_transaction
Modern Rails exposes transaction state through:
Article.current_transaction
You can register an operation:
Article.current_transaction.after_commitdo
SearchIndexer.index(article)
end
This makes a service transaction-aware without requiring it to know whether its caller has opened a transaction.
Rails documents this API as a representation of the current transaction, savepoint, or lack of an active transaction. (Ruby on Rails API)
This is particularly interesting for reusable service objects.
For example:
classPublishArticle
defself.call(article)
article.update!(published:true)
Article.current_transaction.after_commitdo
SearchIndexer.index(article)
end
end
end
Now:
PublishArticle.call(article)
works both:
outside transaction
and:
Article.transactiondo
PublishArticle.call(article)
end
The external action can correctly follow the transaction boundary.
Don’t Rescue StatementInvalid Inside a Transaction
This is one of the most important PostgreSQL-specific transaction rules.
Bad:
User.transactiondo
begin
User.create!(email:"existing@example.com")
rescueActiveRecord::StatementInvalid
# ignore
end
User.create!(email:"new@example.com")
end
A database error such as a unique constraint violation can leave the PostgreSQL transaction in an aborted state.
After that, subsequent SQL statements can fail with an error similar to:
current transaction is aborted,
commands ignored until end of transaction block
Rails explicitly recommends restarting the entire transaction after ActiveRecord::StatementInvalid, rather than continuing within the damaged transaction.
Better:
begin
User.transactiondo
create_user!
create_profile!
end
rescueActiveRecord::RecordNotUnique
# retry or handle outside the transaction
end
The key idea is:
Database failure
↓
Transaction may be unusable
↓
Exit transaction
↓
Handle/retry outside it
This is especially important when building retry logic for concurrency errors.
Transactions Are Not Distributed Transactions
A Rails transaction normally operates on one database connection.
Therefore:
User.transactiondo
user.save!
AuditLog.create!
end
works when those models participate in the same database connection.
But imagine:
Primary DB
User
Analytics DB
AnalyticsEvent
A transaction on the primary database cannot automatically roll back a transaction on another database connection.
Rails explicitly documents that transactions are not distributed across database connections.
This becomes especially important with Rails multiple-database applications.
Multiple Databases: Don’t Assume One Transaction
Imagine:
User.transactiondo
user.update!
AnalyticsEvent.transactiondo
analytics_event.save!
end
end
These are potentially separate database transactions.
You don’t suddenly have:
BEGIN DB1
BEGIN DB2
COMMIT DB1
COMMIT DB2
with a globally atomic guarantee.
Instead, you have two independent database resources.
This is where architectural patterns such as:
transactional outbox
event-driven processing
retries
idempotency
compensating actions
become more appropriate than trying to force a distributed transaction.
Transactions and Background Jobs
Consider:
Order.transactiondo
order.update!(status:"confirmed")
OrderConfirmationJob.perform_later(order.id)
end
This can be dangerous.
Depending on timing, the job could execute before the surrounding transaction has committed.
Then the worker might query:
Order.find(order_id)
and not observe the expected committed state.
Instead:
Order.transactiondo
order.update!(status:"confirmed")
order.after_commitdo
OrderConfirmationJob.perform_later(order.id)
end
end
Or use the appropriate transactional callback mechanisms.
The principle is:
Don’t allow asynchronous consumers to depend on database state that hasn’t committed yet.
Rails’ transaction callbacks are specifically designed for such post-commit work.
Keep Transactions Small
A transaction should generally cover the minimum amount of work necessary.
Avoid:
Order.transactiondo
order.update!
HTTP.get(payment_api)
HTTP.get(shipping_api)
expensive_calculation
sleep(5)
order.update!
end
Now the database transaction stays open while waiting on external systems.
That can mean:
transaction open
↓
database connection occupied
↓
locks potentially held
↓
other requests wait
↓
throughput decreases
A better architecture is often:
external preparation
↓
short DB transaction
↓
commit
↓
after_commit / job
↓
external side effect
Transaction Boundary vs Business Operation
A useful senior-level rule is:
A transaction boundary should normally correspond to a business operation that must be atomic.
For example:
Order.transactiondo
create_order!
reserve_inventory!
record_payment!
end
That’s a meaningful transaction.
But this:
User.transactiondo
user.update!
end
may be unnecessary if you’re only performing one persistence operation.
Remember:
user.update!
already has transactional protection around the persistence operation.
Testing Transaction Behavior
Transactions are especially valuable to test explicitly.
Example:
it"rolls back the order when payment fails"do
expect {
CheckoutService.call(user, cart)
}.toraise_error(PaymentError)
expect(Order.count).toeq(0)
end
Test the business guarantee, not the implementation detail.
Good transaction tests answer questions such as:
Does failed payment rollback the order?
Does failed inventory reservation rollback the payment?
Does an after_commit job run only after successful commit?
Does a nested requires_new operation rollback independently?
A Practical Senior-Level Example
Let’s build a realistic checkout flow.
class CheckoutService
def call(user:, cart:)
order = nil
Order.transaction do
order = Order.create!(
user: user,
total: cart.total,
status: "pending"
)
reserve_inventory!(cart)
Payment.create!(
order: order,
amount: cart.total,
status: "paid"
)
order.update!(status: "confirmed")
ActiveRecord::after_all_transactions_commit do
OrderConfirmationJob.perform_later(order.id)
end
end
order
end
private
def reserve_inventory!(cart)
cart.items.each do |item|
item.product.with_lock do
raise OutOfStock if item.product.stock < item.quantity
item.product.update!(
stock: item.product.stock - item.quantity
)
end
end
end
end
There are several senior-level ideas here.
Atomicity
Order.transaction
ensures the order, payment and inventory changes form one unit.
Concurrency control
with_lock
protects inventory from concurrent updates.
Post-commit processing
ActiveRecord.after_all_transactions_commit
prevents the job from being dispatched before the transaction chain is complete.
This is much closer to production-grade transaction design than simply knowing:
Model.transactiondo
end
Transaction APIs at a Glance
API
Main purpose
Typical usage
Model.transaction
Transaction around a unit of work
Service objects
instance.transaction
Transaction associated with a model instance
Model-centric workflows
ApplicationRecord.transaction
Application-wide transaction boundary
Shared models
transaction(requires_new: true)
Independent nested savepoint
Isolating sub-operations
transaction(isolation: :serializable)
Stronger concurrency guarantees
Highly concurrent workflows
with_lock
Transaction + row lock
Balance/inventory updates
after_commit
Run code after commit
External side effects
after_rollback
React to rollback
Cleanup/recovery logic
transaction.after_commit
Callback attached to a specific transaction
Service/domain workflows
ActiveRecord.after_all_transactions_commit
Run after outermost transaction chain commits
Transaction-aware reusable services
current_transaction.after_commit
Make services transaction-aware
Reusable domain services
Rails provides all of these around the same fundamental transaction system.
How I Decide Which API to Use
As a practical decision tree:
One database operation
Usually:
user.update!
No explicit transaction required.
Several operations must succeed together
Use:
User.transactiondo
...
end
Reusable service may be called inside another transaction
Consider:
transaction(requires_new:true)
when independent rollback semantics are actually required.
Concurrent modification of one row
Use:
record.with_lockdo
...
end
External system must run only after DB success
Use:
after_commit
or a transaction-aware post-commit mechanism.
Multiple database connections
Don’t assume a single transaction protects everything.
But don’t assume your Ruby object has magically reverted every piece of in-memory state to its pre-transaction state. Rails explicitly notes that database rollback doesn’t restore Active Record objects to their original in-memory state.
If you are building AI features into a Rails, Node.js, Python, or any other application, you quickly run into a practical problem:
Which AI model should I use?
OpenAI? Claude? Gemini? DeepSeek? Llama? Mistral?
And what happens when your chosen provider is expensive, rate-limited, unavailable, or simply not the best model for a particular task?
This is where OpenRouter becomes interesting.
OpenRouter provides a unified API for accessing hundreds of AI models through a single interface. It follows an OpenAI-compatible API style, so applications using the OpenAI SDK can often switch to OpenRouter with very little code change. (OpenRouter)
What is OpenRouter?
Think of OpenRouter as an AI gateway/router sitting between your application and multiple LLM providers.
Instead of:
Your Application
|
+----> OpenAI
|
+----> Anthropic
|
+----> Google
|
+----> DeepSeek
you can have:
Your Application
|
v
OpenRouter
|
+----> OpenAI
+----> Anthropic
+----> Google
+----> DeepSeek
+----> Meta
+----> Other providers
Your application talks to one API, while OpenRouter handles access to the underlying models and providers.
It currently exposes hundreds of models through its API, and the available catalog can be queried programmatically. (OpenRouter)
Why would a developer use it?
The biggest advantage isn’t simply “many models.”
The real advantage is reducing coupling to a single AI provider.
Imagine your Rails application has:
MODEL="some-expensive-model"
Six months later you discover that another model:
performs better for your use case
costs less
has better latency
has higher availability
With a direct provider integration, changing providers can involve SDKs, authentication, request formats, response formats and application-specific code.
With OpenRouter, the model is largely a configuration decision:
MODEL="provider/model-name"
That makes experimentation much easier.
Practical Example: OpenAI-Compatible API
One of the most useful features is OpenAI API compatibility.
For example, using the OpenAI Ruby client, the important difference is the base_url:
The exact Ruby client API can vary by gem version, but the architectural idea is simple:
Keep your application code mostly unchanged and change the endpoint/model configuration.
OpenRouter officially documents using the OpenAI SDK with its API by changing the baseURL to the OpenRouter endpoint. (OpenRouter)
Which ruby gem to use?
1. The Recommended Path: The Official openai Gem (Drop-in Compatibility)
# AI assistant - OpenAI
gem "openai", "< 2.0"
Because OpenRouter mirrors OpenAI’s API structure, the easiest and most stable approach is to use the popular official-adjacent openai gem. You simply swap out the base_url and pass your OpenRouter API key.
My Current Rails Implementation is given below (Edited)
While OpenRouter does not maintain an official, first-party SDK exclusively for Ruby, its API is fully OpenAI-compatible. This gives you three simple ways to integrate OpenRouter into a Ruby application
You can test the same prompt against different models without building three separate integrations.
This is particularly useful during development.
For example:
Task: Generate SQL query from natural language
Model A → Good accuracy, expensive
Model B → Very good accuracy, cheaper
Model C → Fast, acceptable accuracy
Instead of making a permanent decision immediately, you can benchmark them.
That’s a much better engineering approach than blindly choosing a model because it is popular.
This is one of the features I find particularly useful for production systems.
Suppose your primary model is temporarily:
Rate limited
↓
Provider outage
↓
Model unavailable
OpenRouter can automatically try another model/provider according to your routing configuration. (OpenRouter)
For example:
models:[
"primary-model",
"fallback-model-1",
"fallback-model-2"
]
If the first model fails, OpenRouter can attempt the next one.
This turns your AI integration from:
Application → One AI Provider
into something closer to:
Application
|
v
OpenRouter
|
+---- Primary
|
+---- Fallback
|
+---- Another fallback
For production applications, that resilience can be more important than simply having access to many models.
Provider Routing
There is another layer that is easy to overlook.
A model may be available through multiple providers.
OpenRouter can route requests between providers and allows developers to influence routing based on things such as provider order, price, throughput and latency. (OpenRouter)
For example, if your application cares primarily about speed, routing can be configured to prefer higher-throughput providers.
If cost is the priority, you can prioritize price.
That means your architecture can move from:
Use Model X
towards:
Use Model X
through the provider that currently makes the most sense
That is a much more interesting abstraction for production AI systems.
What About Cost?
OpenRouter doesn’t magically make every model free.
The underlying model still has its own pricing.
OpenRouter says it passes through provider pricing while providing unified billing and routing. (OpenRouter)
However, OpenRouter also exposes free models.
For example:
openrouter/free
is available as a free-model option, subject to the applicable limits. (OpenRouter)
This is particularly useful when learning or experimenting.
For example, instead of spending money while learning AI API integration:
Rails App
↓
OpenRouter
↓
Free/low-cost model
You can first build the feature, understand the API, streaming, prompts and error handling, and only later move to a more capable paid model.
Important: free does not mean unlimited. OpenRouter documents rate limits for free models, and those limits depend on account/credit conditions. (OpenRouter)
🏗️ A Good Architecture for Rails
For a Rails application, I wouldn’t scatter OpenRouter calls throughout controllers.
Instead, create an abstraction:
class AiClient
def initialize
@client = OpenAI::Client.new(
access_token: ENV["OPENROUTER_API_KEY"],
base_url: "https://openrouter.ai/api/v1"
)
end
def ask(prompt)
@client.chat(
parameters: {
model: ENV.fetch("AI_MODEL"),
messages: [
{ role: "user", content: prompt }
]
}
)
end
end
Then your application does:
response=AiClient.new.ask(
"Summarize this customer feedback"
)
The model becomes configuration:
AI_MODEL=provider/model-name
Now changing the model doesn’t require changing business logic.
That’s the pattern I would recommend for a production Rails application.
Where OpenRouter Makes the Most Sense
I would consider OpenRouter when:
1. You are experimenting with multiple LLMs
You don’t want to build five separate integrations just to compare models.
2. You want provider flexibility
Your application shouldn’t become tightly coupled to one AI company unless there is a strong reason.
3. You need fallback strategies
AI APIs can experience rate limits and provider outages. Model/provider fallback can improve resilience. (OpenRouter)
4. You are cost-conscious
You can compare models and route workloads according to cost/performance requirements.
5. You are building an AI abstraction layer
For example:
Rails Application
|
v
AiClient
|
v
OpenRouter
|
+---+---+---+
| | | |
GPT Claude Gemini DeepSeek
Your business logic doesn’t need to know which provider actually processed the request.
Should You Always Use OpenRouter?
No.
There are situations where going directly to the provider makes more sense.
For example, if your application is deeply dependent on provider-specific features, you may want the official SDK/API directly.
Also, adding another layer means you should evaluate:
latency
provider availability
data/privacy requirements
supported API features
model-specific behavior
operational dependencies
OpenRouter also provides controls around provider selection and data collection, including options such as Zero Data Retention routing where supported, so these requirements should be evaluated rather than assumed. (OpenRouter)
My Take as a Senior Developer
I wouldn’t look at OpenRouter simply as “a website where I can access different AI models.”
The more interesting way to think about it is:
OpenRouter is an abstraction layer between your application and the rapidly changing LLM ecosystem.
The AI world is moving extremely fast.
Today’s best model may not be tomorrow’s best model.
If your application is tightly coupled to:
Application → Provider SDK → One Model
you have created an architectural dependency.
If instead you build:
Application
↓
AI Service / Adapter
↓
OpenRouter
↓
Multiple Models / Providers
you gain considerably more flexibility.
For me, model experimentation, provider independence, automatic fallback and a consistent API are the strongest reasons to consider OpenRouter.
And for someone learning AI development, it is also a practical way to experiment with different models without writing a completely different integration for every provider.
Bottom line: If you’re building AI features today, don’t think only about which model to use. Think about how easily you can change that model tomorrow. OpenRouter is one practical way to design for that flexibility.
Conceptually, that becomes a pgvector nearest-neighbor query using cosine distance. pgvector supports cosine distance through <=>.
3. Test semantic search
You already have three chunks:
Chunk 1
Ruby blocks are chunks of code passed to methods.
Chunk 2
Ruby modules allow code to be organized and reused.
Chunk 3
Ruby classes define objects and their behavior.
Let’s test with a query that doesn’t use the exact wording from the second chunk.
Run:
bin/rails c
Then:
search=Ai::VectorSearchService.new
Now:
results=search.call(
query:"How can I reuse code in Ruby?"
)
Inspect:
results.map(&:content)
You should ideally see the modules chunk near the top:
"Ruby modules allow code to be organized and reused."
That’s our first semantic retrieval.
4. See the ranking
I want you to see why the result was selected.
Ask for the distance:
results.map do |chunk|
{
id: chunk.id,
content: chunk.content,
distance: chunk.neighbor_distance
}
end
Depending on your Neighbor version, the distance accessor may be exposed differently. If neighbor_distance isn’t available, don’t spend time debugging it yet; the returned ordering is the important part for this checkpoint.
A basic test can use a fake embedding service, because we don’t want every test to call the embedding API.
require "test_helper"
class Ai::VectorSearchServiceTest < ActiveSupport::TestCase
test "returns nearest document chunks" do
document = Document.create!(title: "Ruby Guide", source: "test")
document.document_chunks.create!(
content: "Ruby blocks are passed to methods.",
chunk_index: 0,
embedding: Array.new(1024, 0.1)
)
document.document_chunks.create!(
content: "Ruby modules allow code reuse.",
chunk_index: 1,
embedding: Array.new(1024, 0.2)
)
fake_embedding_service = Minitest::Mock.new
fake_embedding_service.expect(
:call,
Array.new(1024, 0.2),
text: "How do I reuse Ruby code?"
)
service = Ai::VectorSearchService.new(
embedding_service: fake_embedding_service
)
results = service.call(
query: "How do I reuse Ruby code?",
limit: 1
)
assert_equal 1, results.size
assert_equal "Ruby modules allow code reuse.", results.first.content
fake_embedding_service.verify
end
end
Because our vectors are artificial, this test is mainly verifying the service’s wiring. For higher-confidence semantic-search tests, we’d later use controlled fixtures or a small integration test.
Next – The Actual RAG Answer
We’re now one step away from having a real RAG feature.
Currently:
Question
↓
VectorSearchService
↓
Relevant chunks
Next we’ll do:
Question
↓
VectorSearchService
↓
Top 5 chunks
↓
PromptBuilder
↓
LLM
↓
Answer grounded in document
We’ll modify Ai::PromptBuilder so it can accept retrieved context and implement:
Ai::RagService
That will be the point where our Rails app goes from “I can search vectors” to “I have built a RAG application.”
Step 14 – Complete RAG in Rails
Now we have reached the final step of the RAG implementation:
class Ai::PromptBuilder
SYSTEM_PROMPT = <<~PROMPT
You are a helpful AI assistant.
Answer questions clearly and concisely.
When document context is provided:
- Use the provided context as the primary source of truth.
- Do not invent information that is not supported by the context.
- If the answer cannot be determined from the context, say that you don't have enough information.
PROMPT
def initialize(conversation:, context: nil)
@conversation = conversation
@context = context
end
def build
messages = [
{
role: "system",
content: SYSTEM_PROMPT.strip
}
]
if @context.present?
messages << {
role: "system",
content: <<~CONTEXT
Use the following document context to answer the user's question:
#{@context}
CONTEXT
}
end
messages.concat(
@conversation.messages
.order(:created_at)
.map do |message|
{
role: message.role,
content: message.content
}
end
)
messages
end
end
result = rag.call(
conversation: conversation,
question: "How can Ruby code be reused?"
)
Then:
putsresult[:content]
You should get an answer based on your stored document chunks.
Error: I didn’t see a question in your message, This model is unavailable for free
When I tested result = rag.call(... I get this error:
Error #2
{url: "https://openrouter.ai/api/v1/chat/completions", status: 404, body: {error: {message: "This model is unavailable for free. The paid version is available now - use this slug instead: minimax/minimax-m3", code: 404}, user_id: "user_xxxxx"}} (OpenAI::Errors::NotFoundError)
Our model fallback does not work. So I used another one as our first model from the models array: "nvidia/nemotron-3-super-120b-a12b:free", Then it works. But I get the following result:
Error #1
puts result
{content: "I'm ready to answer your question, but I didn’t see a question in your message. Could you please provide the question you’d like me to answer based on the given context about Ruby modules, blocks, and classes?", model: "nvidia/nemotron-3-super-120b-a12b:free", input_tokens: 139, output_tokens: 151}
Why is it?
There are actually two separate issues here. The first one is very clear from our RagService result.
1. Why the RAG answer says “I didn’t see a question”
Your code currently does this:
result=rag.call(
conversation:conversation,
question:"How can Ruby code be reused?"
)
But in our Ai::RagService, we only use question for retrieval:
chunks=@vector_search_service.call(
query:question,
limit:limit
)
Then we build the prompt from:
Ai::PromptBuilder.new(
conversation:conversation,
context:context
).build
The question itself is never added to the conversation or prompt.
So the LLM receives approximately:
SYSTEM:
You are a helpful AI assistant...
CONTEXT:
Ruby blocks are...
Ruby modules allow code to be organized and reused.
Ruby classes define...
[previous conversation, if any]
but not:
USER:
How can Ruby code be reused?
Therefore the model is actually doing the right thing when it says:
“I didn’t see a question in your message.”
This is our bug.
2. Fix Ai::RagService
We need to make the current question part of the LLM input.
I actually prefer making the question a first-class argument to PromptBuilder rather than appending it afterward.
So our cleaner final API can become:
Ai::PromptBuilder.new(
conversation:conversation,
context:context,
current_question:question
).build
Then PromptBuilder controls the complete LLM prompt.
We’ll do that after confirming the current fix works.
4. About your fallback problem
You’re also correct that the fallback behavior isn’t happening as expected.
Your error:
This model is unavailable for free.
The paid version is available now...
came back as:
OpenAI::Errors::NotFoundError
That’s HTTP 404.
OpenRouter’s current documentation says its models array should trigger fallback when the primary model returns an error, and when using the OpenAI SDK it should be supplied through extra_body. (OpenRouter)
However, there’s an important detail in our current Ruby SDK usage.
The latest OpenAI Ruby SDK documentation says undocumented request parameters such as OpenRouter’s models extension should be passed using:
request_options: {
extra_body: {
models: [...]
}
}
not simply:
extra_body: {
models: [...]
}
The SDK documents extra_body specifically under request_options. (GitHub)
So our earlier code was likely passing the OpenRouter extension in the wrong place.
5. Fix Ai::Client fallback request
Change this:
response=@client.chat.completions.create(
model:MODELS.first,
messages:messages,
extra_body: {
models:MODELS.drop(1)
}
)
to:
response=@client.chat.completions.create(
model:MODELS.first,
messages:messages,
request_options: {
extra_body: {
models:MODELS.drop(1)
}
}
)
That’s the key fix.
The OpenAI Ruby SDK explicitly documents request_options.extra_body for passing provider-specific/undocumented request parameters.
6. Test the fallback independently
Before retesting RAG, let’s isolate fallback.
Temporarily make:
MODELS= [
"an-invalid-or-unavailable-model",
"nvidia/nemotron-3-super-120b-a12b:free"
].freeze
Then:
client = Ai::Client.new
result = client.chat(
messages: [
{
role: "user",
content: "Why is Node.js commonly used as a backend?"
}
]
)
Then:
putsresult[:content]
putsresult[:model]
We want:
requested primary → fails
fallback → succeeds
result[:model]
=> "nvidia/nemotron-3-super-120b-a12b:free"
If that works, restore your real MODELS.
This is a much better test than testing fallback through the full RAG stack.
* Now Let’s Move On to Our Development.
14.4 See the actual retrieved context
Before trusting the final answer, inspect retrieval independently:
search = Ai::VectorSearchService.new
chunks = search.call(
query: "How can Ruby code be reused?",
limit: 3
)
Then:
chunks.eachdo |chunk|
puts"-----"
putschunk.content
end
You should see something like:
-----
Ruby modules allow code to be organized and reused.
-----
Ruby classes define objects and their behavior.
That’s the crucial RAG mechanism.
The model didn’t search PostgreSQL.
Rails searched PostgreSQL first and gave the model the relevant information.
14.5 Connect RAG to the chat flow
Right now our application uses:
Ai::ChatService
for normal chat.
We can keep that and add a dedicated RAG path.
For example, create an endpoint/action later such as:
DOCUMENT INGESTION
PDF / Document
↓
Text Extraction
↓
Chunking
↓
EmbeddingService
↓
Embedding Model
↓
vector(1024)
↓
DocumentChunk
↓
PostgreSQL + pgvector
14.12 The answer you should memorize
Question:
“Explain how you implemented RAG in Rails.”
You can now say:
“I split documents into chunks and generated embeddings for each chunk. I stored those embeddings in PostgreSQL using pgvector. At query time, I embed the user’s question and perform cosine similarity search to retrieve the most relevant chunks. I then inject those chunks as context into the prompt and send the augmented prompt to the LLM. I keep retrieval, prompt construction and provider communication behind separate Rails services.”
That’s already enough material for a serious Senior Rails + AI int. discussion.
Now, we’ll consider RAG mechanically complete and move to the next major bootcamp topic: AI Agents + Tool Calling, where we’ll turn the assistant from:
Question → Answer
into:
Question
↓
Agent
├── Search docs
├── Search products
├── Find order
└── Execute business action
That is where your Rails business-logic and API design experience becomes especially valuable.
One important correction before we code: the free model list you fetched earlier contains no free embedding model slug. OpenRouter currently lists liquid/lfm2.5-embedding-350m as a free embedding model, producing 1,024-dimensional vectors. OpenRouter’s embeddings API is OpenAI-compatible, so we can use the same Ruby SDK/base URL. (OpenRouter)
That means our existing vector(1536) column is the wrong dimension for the free embedding model we’ll use. We’ll fix that now.
def embed(text:)
response = @client.embeddings.create(
model: EMBEDDING_MODEL,
input: text
)
{
embedding: response.data.first.embedding,
model: response.model,
input_tokens: response.usage&.prompt_tokens
}
end
So your client now has two responsibilities:
chat()
embed()
Both communicate with the same OpenRouter endpoint, but use different models/endpoints. OpenRouter provides an OpenAI-compatible /embeddings API for this. (OpenRouter)
Step 4 – Test the raw embedding request
Open Rails console:
bin/rails c
Then:
client=Ai::Client.new
Now:
result = client.embed(
text: "Ruby on Rails is a web application framework."
)
Inspect:
result.keys
You should get:
[:embedding, :model, :input_tokens]
Now:
result[:embedding].length
You should get:
1024
This is an important RAG checkpoint.
You’ve just proven:
text
↓
embedding model
↓
1024 numbers
Now inspect the first few values:
result[:embedding].first(5)
You’ll see floating-point numbers.
Don’t worry about the actual values. Their position in vector space is what matters.
Step 5 – Create Ai::EmbeddingService
Now we introduce the application-level service.
Create:
app/services/ai/embedding_service.rb
Use:
class Ai::EmbeddingService
def initialize(ai_client: Ai::Client.new)
@ai_client = ai_client
end
def call(text:)
result = @ai_client.embed(text: text)
result[:embedding]
end
end
Why create another service when Ai::Client already has embed?
Because these are different responsibilities:
Ai::Client
How do I communicate with OpenRouter?
Ai::EmbeddingService
How does our application generate an embedding?
That distinction becomes useful once we introduce:
We now have our first actual vector stored in PostgreSQL.
Error: I cannot update embedding vector column with Ruby Array embedding data
I have tested to storing the embedding. But it seems to be Rails does not know / there is a Type mismatch for embedding ruby array data and db vector data type
Question
↓
Query embedding
↓
pgvector
↓
ORDER BY cosine distance
↓
Top K chunks
That is the point where RAG becomes real.
Then we’ll build Ai::VectorSearchService and make the first semantic search against PostgreSQL – the most important practical RAG step after embeddings.
We’ll move quickly, but this time keep each milestone runnable. Since you already have PostgreSQL and a working Rails 8.1 app, pgvector is a natural fit: it stores vectors alongside normal PostgreSQL data and supports cosine similarity plus exact and approximate nearest-neighbor search. (GitHub)
Step 13A – Install and enable pgvector
1. Check your PostgreSQL version
Run:
psql --version
Then check whether the extension is already installed:
bin/rails dbconsole
Inside PostgreSQL:
SELECT extname, extversion
FROM pg_extension
WHERE extname = 'vector';
If you get a row
For example:
vector | 0.8.6
you’re ready.
If you get no rows
You need to install the extension on your PostgreSQL installation.
Since you’re on macOS, if PostgreSQL was installed via Homebrew:
brew install pgvector
The pgvector project currently documents Homebrew installation for PostgreSQL 17/18 formulas. (GitHub)
Then restart PostgreSQL if required by your installation:
brew services restart postgresql@14
Use your actual PostgreSQL version if different.
Step 13B – Enable pgvector in Rails
Once PostgreSQL has the extension available, exit psql:
\q
Generate the migration:
bin/rails generate migration EnablePgvector
Open the migration and use:
class EnablePgvector < ActiveRecord::Migration[8.1]
def change
enable_extension "vector"
end
end
Then:
bin/rails db:migrate
Error: PG::UndefinedFile: ERROR: could not open extension control file "/opt/homebrew/share/postgresql@14/extension/vector.control": No such file or director
This error occurs because the pgvector extension is not installed or cannot be found in the directory of your specific Homebrew-managed PostgreSQL 14 installation.
Do:
# 1. Clone the pgvector repository
cd /tmp
git clone --branch v0.8.6 https://github.com/pgvector/pgvector.git
cd pgvector
# 2. Explicitly point to your PostgreSQL 14 pg_config binary
export PG_CONFIG=/opt/homebrew/opt/postgresql@14/bin/pg_config
# 3. Build and install the extension
make
make install # may need sudo
# Verify the Installation: after the installation completes successfully, check if the vector.control file is present in the target directory
ls /opt/homebrew/share/postgresql@14/extension/vector.control
Verify:
➜ ai_assistant git:(main) rails dbconsole
psql (14.17 (Homebrew))
Type "help" for help.
ai_assistant_development=# SELECT extname, extversion
FROM pg_extension
WHERE extname = 'vector';
extname | extversion
---------+------------
(0 rows)
ai_assistant_development=#
\q
➜ ai_assistant git:(main) ✗ brew services restart postgresql@14
Stopping `postgresql@14`... (might take a while)
==> Successfully stopped `postgresql@14` (label: sh.brew.postgresql@14)
==> Successfully started `postgresql@14` (label: sh.brew.postgresql@14)
➜ ai_assistant git:(main) ✗ rails dbconsole
psql (14.17 (Homebrew))
Type "help" for help.
ai_assistant_development=# SELECT extname, extversion
FROM pg_extension
WHERE extname = 'vector';
extname | extversion
---------+------------
vector | 0.8.6
(1 row)
Chunk text
↓
Embedding API
↓
[0.021, -0.318, ...]
↓
PostgreSQL vector column
We’ll use 1536 dimensions initially, because we’ll use an embedding model that produces 1536-dimensional vectors. The actual dimension must match the embedding model you choose; pgvector requires the declared vector dimension to match stored vectors.
Step 13D – Create Document
Run:
bin/rails g model Document title:string source:string
Then:
bin/rails db:migrate
Open:
app/models/document.rb
Change it to:
class Document < ApplicationRecord
has_many :document_chunks, dependent: :destroy
validates :title, presence: true
end
Step 13E – Create DocumentChunk
Generate it:
bin/rails g model DocumentChunk \
document:references \
content:text \
chunk_index:integer
Then don’t migrate yet.
We need to add the vector column manually because Rails’ generator doesn’t know which embedding dimension we want.
Open the generated migration and make it:
class CreateDocumentChunks < ActiveRecord::Migration[8.1]
def change
create_table :document_chunks do |t|
t.references :document, null: false, foreign_key: true
t.text :content, null: false
t.integer :chunk_index, null: false
t.vector :embedding, limit: 1536
t.timestamps
end
add_index(
:document_chunks,
[:document_id, :chunk_index],
unique: true
)
end
end
Depending on the pgvector Rails integration available in your environment, t.vector may not be recognized. If that happens, we’ll use:
-- create_table(:document_chunks)
bin/rails aborted!
StandardError: An error has occurred, this and all later migrations canceled: (StandardError)
undefined method 'vector' for an instance of ActiveRecord::ConnectionAdapters::PostgreSQL::TableDefinition
Do:
rails g migration addEmbeddingToDocumentChunks
# add
add_column :document_chunks, :embedding, :vector, limit: 1536
# do
rails db:migrate -t
Step 13F – Model association
Open:
app/models/document_chunk.rb
Use:
class DocumentChunk < ApplicationRecord
belongs_to :document
validates :content, presence: true
validates :chunk_index, presence: true
end
Step 13G – Verify the database
Run:
bin/rails dbconsole
Then:
\d document_chunks
You should have:
ai_assistant_development=# \d document_chunks
Table "public.document_chunks"
Column | Type | Collation | Nullable | Default
-------------+--------------------------------+-----------+----------+---------------------------------------------
id | bigint | | not null | nextval('document_chunks_id_seq'::regclass)
document_id | bigint | | not null |
content | text | | not null |
chunk_index | integer | | not null |
created_at | timestamp(6) without time zone | | not null |
updated_at | timestamp(6) without time zone | | not null |
embedding | vector | | |
Indexes:
"document_chunks_pkey" PRIMARY KEY, btree (id)
"index_document_chunks_on_document_id" btree (document_id)
"index_document_chunks_on_document_id_and_chunk_index" UNIQUE, btree (document_id, chunk_index)
Foreign-key constraints:
"fk_rails_99b41ada32" FOREIGN KEY (document_id) REFERENCES documents(id)
And:
SELECT vector_dims(
'[1,2,3]'::vector
);
should return:
3
That proves the extension itself is working.
Exit:
\q
Step 13H – Create your first document manually
Before worrying about PDFs, parsers, Sidekiq, etc., let’s prove the RAG data model.
document.document_chunks.create!(
content: "Ruby blocks are chunks of code passed to methods.",
chunk_index: 0
)
document.document_chunks.create!(
content: "Ruby modules allow code to be organized and reused.",
chunk_index: 1
)
document.document_chunks.create!(
content: "Ruby classes define objects and their behavior.",
chunk_index: 2
)
Check:
document.document_chunks.count
Expected:
3
Step 13I – What we’ve built
Our database is now:
documents
----------------
id
title
source
│
│ 1 → many
▼
document_chunks
----------------
id
document_id
content
chunk_index
embedding
The crucial field is:
embedding
which will eventually contain:
[0.012, -0.883, 0.217, ...]
Int. Checkpoint
You should now be able to explain:
Why don’t we put the embedding on documents?
Because a document is usually too large to embed as one semantic unit.
We split it into chunks and embed each chunk independently:
Document
↓
Chunks
↓
Embeddings
That lets retrieval find the relevant section instead of returning the entire document.
One important design choice
We’re not adding an HNSW index yet.
An HNSW (Hierarchical Navigable Small World) index is a high-speed graph-based algorithm used to find similar items in large collections of high-dimensional data. It is widely used in vector databases for AI tasks like semantic search and recommendation systems.
pgvector supports exact nearest-neighbor search by default, and approximate indexes such as HNSW and IVFFlat become useful as the dataset grows. HNSW generally offers a strong speed/recall tradeoff but costs more memory and has a slower build.
IVFFlat (Inverted File with Flat compression) is a type of database index used to speed up similarity searches for high-dimensional vectors
For our small learning dataset:
exact search first
Once we have real embeddings and enough data:
HNSW index
We’ll deliberately compare both, which makes a good senior-level discussion.
We’ll create an Ai::EmbeddingService, generate a real embedding through our current provider setup, store it in PostgreSQL, and then perform our first semantic similarity search. That will be the point where we can honestly say we’ve built RAG mechanics rather than just knowing the definition.
We should implement OpenRouter model fallbacks. I have received an email that is pointing to exactly the right mechanism.
The important distinction is:
model = primary model
models = ordered fallback models
OpenRouter tries the models in order when the current one errors
With the OpenAI Ruby SDK, OpenRouter’s models extension should be passed through extra_body. (OpenRouter)
Also, our previous openai/gpt-oss-20b:free error is precisely the kind of failure where a fallback chain is useful.
1. Don’t use openrouter/free
Let’s make the model selection explicit.
In Ai::Client:
PRIMARY_MODEL="openai/gpt-oss-20b:free"
FALLBACK_MODELS= [
"some-other-free-model:free",
"another-free-model:free"
].freeze
However, don’t blindly copy model names from an old tutorial, because OpenRouter’s free catalog changes. Its current model listing shows multiple free models and their availability/status. (OpenRouter)
For this reason, let’s first see what free models are currently available to your account/API.
This gives us the currently available zero-price model IDs instead of guessing.
Pick 2–3 general-purpose conversational models.
Avoid things whose purpose is:
moderation
safety classification
reranking
embedding
image generation
Our earlier User Safety: safe response is exactly why.
3. Model fallback implementation
I would not use openrouter/free as our primary model anymore and definitely not nvidia/nemotron-3.5-content-safety, which is why you previously got the safety-classification output.
For our AI Assistant app, let’s use three general-purpose free models and let OpenRouter handle model-level fallback. OpenRouter documents that the models array is tried in order and with the OpenAI SDK it belongs inside extra_body. (OpenRouter)
Our free fallback chain
From the models we actually have available, I’d use:
The reason I’m choosing these is that they’re general instruction/chat models rather than specialized safety, embedding, or multimodal models. We are optimizing for learning reliability, not benchmarking model quality.
I would not use:
nvidia/nemotron-3.5-content-safety:free
because that’s the wrong task.
I would also avoid for this particular chat application:
cohere/north-mini-code:free
because we’re building a general assistant rather than a coding-only assistant.
And we won’t use:
openrouter/free
Change Ai::Client
Let’s simplify the configuration.
class Ai::Client
MODELS = [
"minimax/minimax-m3:free",
"google/gemma-4-31b-it:free",
"nvidia/nemotron-3-super-120b-a12b:free"
].freeze
BASE_URL = "https://openrouter.ai/api/v1"
def initialize
api_key = Rails.application.credentials.dig(
:openrouter,
:api_key
)
raise "OpenRouter API key is missing" if api_key.blank?
@client = OpenAI::Client.new(
api_key: api_key,
base_url: BASE_URL
)
end
def chat(messages:)
response = @client.chat.completions.create(
model: MODELS.first,
extra_body: {
models: MODELS.drop(1)
},
messages: messages
)
{
content: response.choices.first.message.content,
model: response.model,
input_tokens: response.usage&.prompt_tokens,
output_tokens: response.usage&.completion_tokens
}
rescue OpenAI::Errors::RateLimitError => e
raise Ai::RateLimitError, e.message
rescue OpenAI::Errors::APITimeoutError => e
raise Ai::TimeoutError, e.message
rescue OpenAI::Errors::APIConnectionError => e
raise Ai::ProviderError, e.message
rescue OpenAI::Errors::APIStatusError => e
raise Ai::ProviderError, e.message
end
end
OpenRouter then tries the models in order if the preceding model can’t serve the request. (OpenRouter)
Why model plus models?
This is worth understanding:
model:MODELS.first
is the primary model.
extra_body: {
models:MODELS.drop(1)
}
are the fallbacks.
So:
M3
↓ unavailable
Gemma
↓ unavailable
Nemotron
If the request succeeds using Gemma, response.model tells us which model actually served the request. OpenRouter documents that the response’s model identifies the model used for the successful run. (OpenRouter)
Test it now
Start:
bin/rails c
Then:
client=Ai::Client.new
And:
result=client.chat(
messages: [
{
role:"user",
content:"Why Node.js as a backend?"
}
]
)
=>
{content:
"# Why Node.js as a Backend?\n\nNode.js has become one of the most popular choices for backend development for several compelling reasons:\n\n## 1. **JavaScript Everywhere**\n- Use the same language (JavaScript) on both frontend and backend\n- Easier to share code between client and server\n- Single language for full-stack development reduces context switching\n\n## 2. **Non-Blocking, Event-Driven Architecture**\n- Built on Google's V8 JavaScript engine\n- Handles thousands of concurrent connections with a single thread\n- Ideal for:...skipping...
=>
> puts result[:model]
minimax/minimax-m3:free
=> nil
Then:
putsresult[:content]
putsresult[:model]
You should now get an actual conversational answer.
Run it several times if you want to observe which model is serving your requests.
And this connects directly to our AiRequest
This is why we built the observability table earlier.
Imagine:
Requested:
minimax/m3
Actual:
google/gemma-4-31b-it
Our admin dashboard should eventually show:
Requested Model minimax/minimax-m3:free
Actual Model google/gemma-4-31b-it:free
Status success
That’s a genuinely useful production metric.
OpenRouter documents that, when using the OpenAI SDK, its models parameter is passed through extra_body. (OpenRouter)
OpenRouter says fallback can happen for provider downtime, rate limiting, moderation refusal and context-length errors, among other errors. (OpenRouter)
4. One thing we should NOT do
Don’t implement this:
begin
call_model_a
rescue
call_model_b
rescue
call_model_c
end
unless you have a very specific reason.
OpenRouter already provides model-level failover and doing the fallback manually would mean:
Your Rails app
↓
request A
↓
failure
↓
request B
while OpenRouter can perform this routing itself.
The provider also knows its own availability and provider-level routing state better than our Rails application does.
So:
Let OpenRouter handle model fallback; let Rails handle application-level error handling.
That’s a clean separation of responsibilities. (OpenRouter)
We have enough practical experience with SSE right now. We don’t need to perfect the transport layer, lets move on to improve our production error handling architecture.
Step 12 – Production Hardening of the AI Integration
class Ai::Error < StandardError
end
class Ai::ProviderError < Ai::Error
end
class Ai::RateLimitError < Ai::Error
end
class Ai::TimeoutError < Ai::Error
end
This gives our application its own error vocabulary instead of exposing SDK/provider exceptions everywhere.
The exact exception classes can depend on the SDK/version, so inspect the exception raised by your installed openai gem rather than blindly copying provider-specific classes.
“I distinguish transient failures from permanent failures. For transient failures, I use a bounded number of retries with exponential (delay: 1,2,4,8,16 seconds) backoff.”
Step 13: Add AI Observability with admin Dashboard
Instead of merely saying we support observability, let’s build an actual AI Admin / Observability dashboard into the app. This will make the project much stronger because you can demonstrate that we thought beyond “call the LLM.”
We will track:
AI Request
├── provider
├── model
├── operation
├── status
├── conversation
├── message
├── input tokens
├── output tokens
├── estimated cost
├── latency
├── started/completed timestamps
├── retry count
├── HTTP status
├── error class
├── error message
├── request ID
├── streamed?
└── metadata
And the admin UI will have:
/admin/ai_requests
AI Observability
-------------------------------------------------
Total Requests 127
Successful 119
Failed 8
Total Input Tokens 45,230
Total Output Tokens 18,921
Avg Latency 2.34 sec
Estimated Cost $0.00 / N/A
-------------------------------------------------
Recent AI Requests
-------------------------------------------------
Time | Model | Status | Tokens | Latency | Error
-------------------------------------------------
...
Then clicking a request gives the complete details.
Step 12A – Create AiRequest
We’ll call the model AiRequest.
This is not the AI message itself.
Remember:
Message
↓
What the user/assistant said
AiRequest
↓
What happened while talking to the LLM
namespace :admin do
resources :ai_requests, only: %i[index show]
end
So your routes become something like:
Rails.application.routes.draw do
resources :conversations, only: [:create, :show] do
resources :messages, only: [:create]
end
namespace :admin do
resources :ai_requests, only: %i[index show]
end
root "conversations#new"
end
Check:
bin/rails routes | grep ai_requests
You should get:
/admin/ai_requests
/admin/ai_requests/:id
Step 12I – Admin Controller
Open:
app/controllers/admin/ai_requests_controller.rb
Use:
class Admin::AiRequestsController < ApplicationController
before_action :authenticate_admin!
def index
@ai_requests = AiRequest
.includes(:conversation, :message)
.recent
.limit(100)
@total_requests = AiRequest.count
@successful_requests =
AiRequest.successful.count
@failed_requests =
AiRequest.failed_requests.count
@total_input_tokens =
AiRequest.sum(:input_tokens)
@total_output_tokens =
AiRequest.sum(:output_tokens)
@average_latency =
AiRequest.where.not(latency_ms: nil).average(:latency_ms)
@estimated_cost =
AiRequest.sum(:estimated_cost)
end
def show
@ai_request = AiRequest.includes(
:conversation,
:message
).find(params[:id])
end
private
def authenticate_admin!
authenticate_or_request_with_http_basic("AI Admin") do |username, password|
username == Rails.application.credentials.dig(:admin, :username) &&
password == Rails.application.credentials.dig(:admin, :password)
end
end
end
This means the admin dashboard isn’t publicly accessible.
here only to demonstrate recording unexpected failures.
In the final production version, we’ll distinguish:
timeout
rate limit
provider error
invalid response
unexpected application bug
and map them to the proper AiRequest.status.
That’s coming immediately after this.
Why this dashboard is worth having
You now have a tangible answer to questions like:
How would you monitor an AI application?
You can say:
“I record each AI invocation separately from the conversation message itself. I track provider, model, status, latency, token consumption, retries, HTTP status and error information, then expose that through an internal observability dashboard.”
Then show page:
/admin/ai_requests
That’s much stronger than saying:
“I would use logging.”
One thing I deliberately did NOT add
I don’t recommend storing the complete prompt by default in AiRequest.
Why?
Because prompts can contain:
PII
customer data
confidential company information
secrets
Instead we can later store safe metadata such as:
{
"message_count":8,
"prompt_tokens":1200,
"temperature":0.2
}
and keep sensitive content under the normal conversation access controls.
Issue 1:Fix AI Response: User Safety
Currently when I tested I get the AI Response like: User Safety: safeResponse Safety: safe
This is a model-selection problem, not a Rails problem.
The response:
User Safety: safeResponse Safety: safe
is characteristic of a content-safety/guardrail model, not a normal conversational model. OpenRouter currently lists Nemotron 3.5 Content Safety (free) as a moderation model whose intended output is exactly safety classifications such as User Safety and Response Safety. (OpenRouter)
Because we’re using:
MODEL="openrouter/free"
OpenRouter is free to route that request to an available free model. The free-model router is explicitly designed to select among available free models, so you shouldn’t use it when you need a stable application behavior. (OpenRouter)
Fix: choose an actual chat model
For our course, let’s use a specific free conversational model instead of:
MODEL="openrouter/free"
A good current option is:
MODEL="openai/gpt-oss-20b:free"
OpenRouter lists free models separately, including general-purpose models; the exact free catalog changes over time.
Change Ai::Client
Open:
app/services/ai/client.rb
Change:
MODEL="openrouter/free"
to:
MODEL="openai/gpt-oss-20b:free"
Then test:
bin/rails c
client=Ai::Client.new
result=client.chat(
messages: [
{
role:"user",
content:"Why Node.js as a backend?"
}
]
)
putsresult[:content]
We should now get an actual explanatory answer rather than the safety classification.
Why I want a specific model for our project
This is actually a valuable AI engineering lesson.
Current approach
Ai::Client
↓
openrouter/free
↓
??? model
The model can change depending on routing.
Better application architecture
Ai::Client
↓
specific model
↓
predictable behavior
For production systems, model choice should generally be deliberate rather than an accidental consequence of a router.
The openrouter/free router is useful for experimentation, but for our course we’ll use an explicit free model so our behavior stays understandable. OpenRouter itself recommends openrouter/free as a convenient way to sample available free models, which is precisely why it shouldn’t be treated as a fixed model identity.
One more thing: our RAG work needs an embedding model
Don’t use the chat model for embeddings.
We’ll have:
Chat:
openai/gpt-oss-20b:free
Embeddings:
separate embedding model
OpenRouter currently lists free embedding models as well, including NVIDIA’s Nemotron 3 Embed 1B, which is specifically intended for retrieval/RAG. (OpenRouter)
We’ll choose the embedding model separately when we implement Ai::EmbeddingService.
For now
Make this one-line change:
MODEL="openai/gpt-oss-20b:free"
After that, we’ll continue with Step 13 – generating embeddings and storing the first real vector in document_chunks.
Issue 2: OpenAI::Errors::NotFoundError
Our server Log:
OpenAI::Errors::NotFoundError ({url: "https://openrouter.ai/api/v1/chat/completions", status: 404, body: {error: {message: "This model is unavailable for free. The paid version is available now - use this slug instead: openai/gpt-oss-20b", code: 404}, user_id: ...
Since we’re using the openai Ruby SDK, our rescue layer should use OpenAI::Errors::*, not Faraday exceptions. The SDK maps HTTP status codes such as 400, 401, 403, 404, 409, 422, 429 and 500+ into its own typed exceptions, and it has separate APIConnectionError / APITimeoutError classes. (https://github.com/openai/openai-ruby/blob/main/lib/openai/errors.rb)
Also, our 404 message tells us something important:
OpenRouter’s current free catalog does include openai/gpt-oss-20b:free, but free endpoints can change availability. (OpenRouter)
Our earlier 404 specifically said that the endpoint was unavailable for free at that moment and suggested the paid slug. Since OpenRouter currently lists the :free variant as free, this looks like provider/availability inconsistency, not that our slug was fundamentally wrong. OpenRouter also notes that free variants are rate-limited and availability can vary. (OpenRouter)
1. Fix the model
Let’s use the explicit free model again:
MODEL="openai/gpt-oss-20b:free"
OpenRouter currently lists that exact slug as free with zero input/output pricing. (OpenRouter)
If that endpoint temporarily fails, we can switch to another currently listed free model rather than using openrouter/free.
2. Fix Ai::Client error handling
Also change our Ai::Client chat rescues from: Faraday::TooManyRequestsError
Faraday::TimeoutError
Faraday::Error
to: similar to: OpenAI::Errors::NotFoundError etc,
check: https://github.com/openai/openai-ruby/blob/main/lib/openai/errors.rb
Let’s use the actual SDK error hierarchy.
The important classes are:
OpenAI::Errors::BadRequestError
OpenAI::Errors::AuthenticationError
OpenAI::Errors::PermissionDeniedError
OpenAI::Errors::NotFoundError
OpenAI::Errors::ConflictError
OpenAI::Errors::UnprocessableEntityError
OpenAI::Errors::RateLimitError
OpenAI::Errors::InternalServerError
OpenAI::Errors::APIConnectionError
OpenAI::Errors::APITimeoutError
The current SDK maps HTTP 404 → NotFoundError, 429 → RateLimitError, and 500+ → InternalServerError. (GitHub)
So replace our old Faraday rescues entirely.
app/services/ai/client.rb
Use:
class Ai::Client
MODEL = "openai/gpt-oss-20b:free"
BASE_URL = "https://openrouter.ai/api/v1"
def initialize
api_key = Rails.application.credentials.dig(:openrouter, :api_key)
raise "OpenRouter API key is missing" if api_key.blank?
@client = OpenAI::Client.new(
api_key: api_key,
base_url: BASE_URL
)
end
def chat(messages:)
response = @client.chat.completions.create(
model: MODEL,
messages: messages
)
{
content: response.choices.first.message.content,
model: response.model,
input_tokens: response.usage&.prompt_tokens,
output_tokens: response.usage&.completion_tokens
}
rescue OpenAI::Errors::RateLimitError => e
raise Ai::RateLimitError, e.message
rescue OpenAI::Errors::APITimeoutError => e
raise Ai::TimeoutError, e.message
rescue OpenAI::Errors::APIConnectionError => e
raise Ai::ProviderError, e.message
rescue OpenAI::Errors::BadRequestError,
OpenAI::Errors::AuthenticationError,
OpenAI::Errors::PermissionDeniedError,
OpenAI::Errors::NotFoundError,
OpenAI::Errors::ConflictError,
OpenAI::Errors::UnprocessableEntityError,
OpenAI::Errors::InternalServerError,
OpenAI::Errors::APIStatusError => e
raise Ai::ProviderError, e.message
end
end
The specific NotFoundError you just encountered will therefore be caught here:
Now let’s implement streaming. OpenRouter supports Server-Sent Events (SSE) when stream: true, and the current Ruby SDK exposes Chat Completions streaming through stream_raw; its higher-level stream helper is not available in every released SDK version. (OpenRouter)
We’ll keep the implementation practical and compatible with the SDK behavior you’re using.
Step 9 – Stream the AI response
What changes?
Currently:
Browser
↓
POST
↓
Rails waits for entire LLM response
↓
redirect
Ai::Client.new.stream_chat(messages: messages) do |delta|
print delta
$stdout.flush
end
You should see the answer appearing progressively:
Ruby is a programming language...
instead of getting the entire answer at once.
Why $stdout.flush?
Ruby can buffer stdout. Flushing makes each chunk visible immediately in the console.
9.2 Now expose streaming from Rails
Instead of making MessagesController#create wait for the completed response, we’ll create a streaming endpoint.
Open:
config/routes.rb
Add:
resources :conversations, only: [:create, :show] do
resources :messages, only: [:create]
end
get "/conversations/:conversation_id/messages/stream",
to: "messages#stream",
as: :conversation_messages_stream
9.3 Add the streaming controller action
Open:
app/controllers/messages_controller.rb
Add:
includeActionController::Live
and:
def stream
conversation = Conversation.find(params[:conversation_id])
response.headers["Content-Type"] = "text/event-stream"
response.headers["Cache-Control"] = "no-cache"
response.headers["X-Accel-Buffering"] = "no"
sse = SSE.new(response.stream)
messages = Ai::PromptBuilder
.new(conversation: conversation)
.build
content = +""
begin
Ai::Client.new.stream_chat(messages: messages) do |delta|
next if delta.blank?
content << delta
sse.write(
{ content: delta },
event: "message"
)
end
sse.write(
{ done: true },
event: "done"
)
ensure
sse.close
response.stream.close
end
end
But Rails doesn’t provide SSE automatically.
Add:
includeActionController::Live
and use Rails’ ActionController::Live::SSE if available in our Rails 8.1 setup, or otherwise we can use the standard SSE format directly. Rails 8.1’s Live controller infrastructure is the relevant mechanism here.
To avoid another dependency, let’s actually use the raw SSE format ourselves.
For our application, we’ll eventually use a Stimulus controller rather than inline JavaScript.
9.7 Don’t spend time styling this
Our immediate objective is proving:
LLM → SSE → Browser
Once you can see the response arriving incrementally, we’ve achieved the important part.
9.8 Commit
Once Ruby streaming works:
git add app/services/ai/client.rb
git commit -m"feat: stream LLM responses"
git push
Then we’ll wire the browser properly.
Int. knowledge from this step
You should now be able to explain:
What is SSE?
A persistent HTTP connection where the server pushes events to the client.
Why use it for AI?
Because LLM output naturally arrives incrementally, and streaming improves perceived latency.
Why not Action Cable?
WebSockets are bidirectional; SSE is simpler when the server primarily needs to push generated output to the browser.
Where does the LLM stream end?
At the Rails server, which consumes the provider’s SSE stream and forwards its own stream to the browser.
OpenRouter documents its AI streaming as SSE, while the Ruby SDK provides streaming chat-completion chunks through stream_raw.
Next step
Since ActionController::Live::SSE exists in Rails 8.1, let’s test the controller before committing.
One important point first: don’t test this through bin/rails server with WEBrick. Rails documents that WEBrick buffers responses, so Live streaming won’t behave correctly. Use our normal Puma server instead. (Ruby on Rails Guides)
1. First verify the route
Run:
bin/rails routes | grep stream
You should see our route, something like:
conversation_messages_stream
GET /conversations/:conversation_id/messages/stream
Then get a conversation ID:
bin/rails c
Conversation.last.id
For example:
1
Exit:
exit
2. Test with curl first
This is the easiest way to prove that the Rails endpoint is actually streaming.
But I prefer curl -N for the first test because the browser doesn’t give you a very useful raw view of SSE events.
Rails’ documentation uses essentially this same pattern – writing to response.stream periodically and closing the stream in ensure. (Ruby on Rails Guides)
One architectural correction before we commit
Don’t commit our current streaming implementation yet.
That is the version worth keeping in our portfolio and discussing in an int. scenario. Rails requires the response headers to be set before the first stream write and requires the stream to be closed when finished. (Ruby on Rails API)
the next step will be to connect the actual user message → streaming endpoint → browser UI rather than having a standalone stream endpoint.
Debug:ActionController::Live::ClientDisconnected – 500 Internal Server Error
Yes – very likely from our rescue behavior, but the deeper issue is that ActionController::Live::ClientDisconnected is not the same exception as IOError in Rails 8.1.
does not necessarily catch the exception you’re seeing.
Why the 500 appears
Our stream is working, then eventually the client closes the connection – or example:
browser finishes and closes the SSE connection
EventSource.close() is called
browser navigates/reloads
user closes the tab
network connection disappears
Rails detects that the client is gone while processing the Live response and raises:
ActionController::Live::ClientDisconnected
Rails’ Live processing happens in a separate thread, and once the response has been committed Rails handles exceptions differently from a normal controller request. (Ruby on Rails API)
Don’t interpret ClientDisconnected as an application failure.
It’s closer to:
Rails: "I'm streaming."
Browser: "I'm no longer listening."
Rails: "Okay."
For SSE, that’s a normal lifecycle event.
Why you’re seeing Completed 500
This is the part that initially looks strange.
With ActionController::Live, Rails starts processing the action in a separate thread. When an exception occurs after the response has already been committed/started streaming, Rails can’t behave like a normal controller request with a fresh 500 response. Rails’ Live implementation has special handling for committed responses and closes the stream after logging the exception.
So our log can be misleading:
Completed 500 Internal Server Error
ActionController::Live::ClientDisconnected
even though:
the user received the streamed answer successfully,
the connection closed normally afterward.
That’s exactly why I would treat ClientDisconnected separately from genuine application errors.
Fix Streaming issue: After entering the question the page get stuck
Our server logs are as follows:
Started POST "/conversations"
Processing by ConversationsController#create as TURBO_STREAM
Parameters: {"...., "conversation" => {"title" => "Tell me about Node js"}, "commit" => "Start conversation"}
Then goes to
Started GET "/conversations/13"
Processing by ConversationsController#show as TURBO_STREAM
Parameters: {"id" => "13"}
Rendering conversations/show.html.erb
200 OK
Then the page stuck on the Form: form_with in the conversations/show.html.erb. the div streaming-response is blank and it is not going to messages controller#stream or I think it is not executing the chat_controller.js also.
Yes – the server logs strongly suggest the problem is before MessagesController#stream. The conversation page loads correctly, then the form submission isn’t triggering our Stimulus chat#submit.
The likely culprit is Turbo + Stimulus registration, not SSE.
Depending on the Rails 8 application template/setup, your index.js may use automatic controller loading instead. The important thing is that chat_controller.js is being registered under:
chat
3. Check application.js
Open:
app/javascript/application.js
You should have the normal Rails setup, typically something along the lines of:
import"@hotwired/turbo-rails"
import"controllers"
The important line is:
import"controllers"
Without it, your Stimulus controllers won’t be registered.