Streaming through your server

If the answer shows up all at once at the end, something between PHP and the browser is holding it back.

Answers are sent to the browser piece by piece, as they are written. If they show up all at once at the end instead, something between PHP and the browser is collecting them first: your web server, PHP-FPM’s connection to it, a compression module, a reverse proxy or a CDN.

01 · What the client already does

For the chat’s answers, the client empties and turns off PHP’s own output buffers and output compression, sends each piece the moment it has it, and marks the answer as a stream nobody should hold back or rewrite (Content-Type: text/event-stream, Cache-Control: no-cache, no-transform, X-Accel-Buffering: no). Nothing in php.ini needs changing.

What it cannot change from PHP is the configuration of the servers in front of it. That part is below.

02 · Apache with PHP-FPM

When Apache hands PHP to PHP-FPM through mod_proxy_fcgi, it collects PHP’s output before sending it on, unless the connection is told to flush. Add this to the virtual host of your site:

<Proxy "fcgi://localhost">
    ProxySet flushpackets=on
</Proxy>

The name in <Proxy> must be the same as the fcgi:// part of the line that sends PHP to PHP-FPM. With this SetHandler, for example, it is fcgi://localhost:

<FilesMatch "\.php$">
    SetHandler "proxy:unix:/run/php/php8.3-fpm.sock|fcgi://localhost"
</FilesMatch>

PHP-FPM must not buffer or compress the answer either. In the php.ini of PHP-FPM (for example /etc/php/8.3/fpm/php.ini):

output_buffering = Off
zlib.output_compression = Off

Then check the configuration, reload Apache and restart PHP-FPM (with your PHP version):

sudo apachectl configtest && sudo systemctl reload apache2
sudo systemctl restart php8.3-fpm

On Red Hat, Rocky or AlmaLinux the service is httpd.

03 · Apache compression

A compressor has to collect text before it can compress it. The .htaccess of the chat already turns mod_deflate off for the answers:

<IfModule mod_setenvif.c>
    SetEnvIf Request_URI "/chat$" no-gzip=1 dont-vary=1
</IfModule>

If your Apache does not read .htaccess in that folder, put the same line in the virtual host. With mod_php instead of PHP-FPM, there is no PHP-FPM connection to set up.

04 · nginx

nginx honours the X-Accel-Buffering: no the client sends. To be explicit, and to keep compression away from the answers, the PHP location of the chat carries two more lines, as on Installing:

location = /opensolr-chat/index.php {
    include fastcgi_params;
    fastcgi_param SCRIPT_FILENAME $document_root/opensolr-chat/index.php;
    fastcgi_pass unix:/run/php/php8.3-fpm.sock;
    fastcgi_buffering off;
    gzip off;
}

When nginx is a reverse proxy in front of another web server, the location that passes the chat on needs the same, for the proxy, and must pass on the visitor’s host name, which the chat checks:

location /opensolr-chat/ {
    proxy_pass http://127.0.0.1:8080;
    proxy_set_header Host $host;
    proxy_buffering off;
    gzip off;
}

Then sudo nginx -t && sudo systemctl reload nginx. If your configuration sets fastcgi_ignore_headers or proxy_ignore_headers with X-Accel-Buffering, the lines above are the only thing that turns buffering off.

05 · Other proxies and CDNs

A load balancer, a caching proxy or a CDN in front of your site must pass /opensolr-chat/chat straight through: no buffering, no compression, no caching. Look in its documentation for streamed responses or Server-Sent Events. When the proxy changes the visitor’s address, also set trusted_proxies, as on Installing.

06 · How to tell it works
  • In the chat: ask a question. The line that says what the assistant is doing should change while you watch, and the words should start appearing within seconds and keep coming.
  • Buffered: nothing moves, then the whole answer appears at once. With a long answer and nothing at all for 45 seconds, the chat gives up with The answer stopped coming. Please try again.

From a shell, with the captcha off (with the captcha on, the answer is a single line asking for the captcha):

curl -N https://your-site/opensolr-chat/chat \
  -H 'Content-Type: application/json' \
  -H 'Origin: https://your-site' \
  --data '{"conversation":"curl-test-1","messages":[{"role":"user","content":"What is this site about?"}]}'

Streamed, the data: lines arrive one by one over several seconds. Buffered, they all arrive together at the end. The question counts toward the limits of your own IP address and uses AI requests like any other.

The Opensolr Chat Bot Client is open source and MIT licensed. Questions about your Opensolr account, index or plan go to opensolr.com/contact; questions about the client itself belong on GitHub.

Opensolr Chat Bot Documentation