Portainer Templates logo

Portainer Templates

TEI Embeddings

Stack

LLM Infrastructure

TEI - Hugging Face's Rust embeddings server: small models like BAAI/bge-small-en-v1.5 (384-dim) serve real RAG workloads from CPU in ~512 MB, over both TEI's native API and an OpenAI-compatible /v1/embeddings endpoint.

Image details

Architecture: amd64
Image size: 228 MB
User: huggingface

Source details

Stars: 5k
Forks: 429
Language: Rust
License: Apache-2.0
Updated: 1 day ago

Configuration

Type
Compose
Platform
linux
Image
ghcr.io/huggingface/text-embeddings-inference:cpu-1.9
Ports
5000:5000
Volumes
/data : modelcache
Env vars
PORT=5000MODEL_ID=${MODEL_ID:-BAAI/bge-small-en-v1.5}HUGGINGFACE_HUB_CACHE=/data
Restart
unless-stopped
Source

Standalone Install

Select an install method, to see config/commands for deploying TEI Embeddings

Installation method

Install on Portainer

Import all app templates into your Portainer instance, for easy 1-click deploys

  1. Ensure both Docker and Portainer are installed, and up-to-date
  2. Log into your Portainer web UI
  3. Under Settings → App Templates, paste the below URL
  4. Head to Home → App Templates, and the list of apps will show up
  5. Select TEI Embeddings, fill in any config options, and hit Deploy

Template Import URL

https://raw.githubusercontent.com/Lissy93/portainer-templates/main/templates.json
Show Me demo
Original stackfile

The compose file this template deploys, straight from its repo:

name: tei

services:
  tei:
    image: ghcr.io/huggingface/text-embeddings-inference:cpu-1.9
    restart: unless-stopped
    environment:
      PORT: "5000"
      MODEL_ID: ${MODEL_ID:-BAAI/bge-small-en-v1.5}
      HUGGINGFACE_HUB_CACHE: /data
    ports:
      - "5000:5000"
    volumes:
      - modelcache:/data

volumes:
  modelcache:

Or deploy it directly from the source:

git clone https://github.com/deployable-sh/stacks
cd stacks
docker compose -f tei/compose.yaml up -d

More install options in our documentation, or see huggingface/text-embeddings-inference for app-specific guidance.

Text Embeddings Inference

GitHub Repo stars Swagger API documentation
A blazing fast inference solution for text embeddings models.
Benchmark for BAAI/bge-base-en-v1.5 on an NVIDIA A10 with a sequence length of 512 tokens:


Table of contents

- [Supported Models](#supported-models)
- [Docker](#docker)
- [Docker Images](#docker-images)
- [API Documentation](#api-documentation)
- [Using a private or gated model](#using-a-private-or-gated-model)
- [Air gapped deployment](#air-gapped-deployment)
- [Using Re-rankers models](#using-re-rankers-models)
- [Using Sequence Classification models](#using-sequence-classification-models)
- [Using SPLADE pooling](#using-splade-pooling)
- [Distributed Tracing](#distributed-tracing)
- [gRPC](#grpc)
- [Apple Silicon (Homebrew)](#apple-silicon-homebrew)
- [ARM64 / aarch64](#arm64--aarch64)

Text Embeddings Inference (TEI) is a toolkit for deploying and serving open source text embeddings and sequence classification models. TEI enables high-performance extraction for the most popular models, including FlagEmbedding, Ember, GTE and E5. TEI implements many features such as:
  • No model graph compilation step
  • Metal support for local execution on Macs
  • Small docker images and fast boot times. Get ready for true serverless!
  • Token based dynamic batching
  • Optimized transformers code for inference using Flash Attention,
Candle and cuBLASLt
  • Safetensors weight loading
  • ONNX weight loading
  • Production ready (distributed tracing with Open Telemetry, Prometheus metrics)

Get Started

Supported Models

Text Embeddings

Text Embeddings Inference currently supports Nomic, BERT, CamemBERT, XLM-RoBERTa models with absolute positions, JinaBERT model with Alibi positions and Mistral, Alibaba GTE, Qwen2 models with Rope positions, MPNet, ModernBERT, Qwen3, and Gemma3.
Below are some examples of the currently supported models:
MTEB RankModel SizeModel TypeModel ID
27.57B (Very Expensive)Qwen3Qwen/Qwen3-Embedding-8B
34.02B (Very Expensive)Qwen3Qwen/Qwen3-Embedding-4B
4509MQwen3Qwen/Qwen3-Embedding-0.6B
67.61B (Very Expensive)Qwen2Alibaba-NLP/gte-Qwen2-7B-instruct
7560MXLM-RoBERTaintfloat/multilingual-e5-large-instruct
8308MGemma3google/embeddinggemma-300m (gated)
151.78B (Expensive)Qwen2Alibaba-NLP/gte-Qwen2-1.5B-instruct
187.11B (Very Expensive)MistralSalesforce/SFR-Embedding-2R
35568MXLM-RoBERTaSnowflake/snowflake-arctic-embed-l-v2.0
41305MAlibaba GTESnowflake/snowflake-arctic-embed-m-v2.0
52335MBERTWhereIsAI/UAE-Large-V1
58137MNomicBERTnomic-ai/nomic-embed-text-v1
79137MNomicBERTnomic-ai/nomic-embed-text-v1.5
103109MMPNetsentence-transformers/all-mpnet-base-v2
N/A475M-A305MNomicBERTnomic-ai/nomic-embed-text-v2-moe
N/A434MAlibaba GTEAlibaba-NLP/gte-large-en-v1.5
N/A396MModernBERTanswerdotai/ModernBERT-large
N/A340MQwen3voyageai/voyage-4-nano
N/A137MJinaBERTjinaai/jina-embeddings-v2-base-en
N/A137MJinaBERTjinaai/jina-embeddings-v2-base-code

To explore the list of best performing text embeddings models, visit the Massive Text Embedding Benchmark (MTEB) Leaderboard.

Sequence Classification and Re-Ranking

Text Embeddings Inference currently supports CamemBERT, and XLM-RoBERTa Sequence Classification models with absolute positions.
Below are some examples of the currently supported models:
TaskModel TypeModel ID
Re-RankingXLM-RoBERTaBAAI/bge-reranker-large
Re-RankingXLM-RoBERTaBAAI/bge-reranker-base
Re-RankingGTEAlibaba-NLP/gte-multilingual-reranker-base
Re-RankingModernBertAlibaba-NLP/gte-reranker-modernbert-base
Sentiment AnalysisRoBERTaSamLowe/roberta-base-goemotions

Docker

model=Qwen/Qwen3-Embedding-0.6B
volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run

docker run --gpus all -p 8080:80 -v $volume:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 --model-id $model

And then you can make requests like
curl 127.0.0.1:8080/embed \
    -X POST \
    -d '{"inputs":"What is Deep Learning?"}' \
    -H 'Content-Type: application/json'

Note: To use GPUs, you need to install the NVIDIA Container Toolkit. NVIDIA drivers on your machine need to be compatible with CUDA version 12.2 or higher.
To see all options to serve your models:
$ text-embeddings-router --help
Text Embedding Webserver

Usage: text-embeddings-router [OPTIONS] --model-id <MODEL_ID>

Options:
      --model-id <MODEL_ID>
          The Hugging Face model ID, can be any model listed on <https://huggingface.co/models> with the `text-embeddings-inference` tag (meaning it's compatible with Text Embeddings Inference).

          Alternatively, the specified ID can also be a path to a local directory containing the necessary model files saved by the `save_pretrained(...)` methods of either Transformers or Sentence Transformers.

          [env: MODEL_ID=]

      --revision <REVISION>
          The actual revision of the model if you're referring to a model on the hub. You can use a specific commit id or a branch like `refs/pr/2`

          [env: REVISION=]

      --tokenization-workers <TOKENIZATION_WORKERS>
          Optionally control the number of tokenizer workers used for payload tokenization, validation and truncation. Default to the number of CPU cores on the machine

          [env: TOKENIZATION_WORKERS=]

      --dtype <DTYPE>
          The dtype to be forced upon the model

          [env: DTYPE=]
          [possible values: float16, float32]

      --served-model-name <SERVED_MODEL_NAME>
          The name of the model that is being served. If not specified, defaults to `--model-id`. It is only used for the OpenAI-compatible endpoints via HTTP

          [env: SERVED_MODEL_NAME=]

      --pooling <POOLING>
          Optionally control the pooling method for embedding models.

          If `pooling` is not set, the pooling configuration will be parsed from the model `1_Pooling/config.json` configuration.

          If `pooling` is set, it will override the model pooling configuration

          [env: POOLING=]

          Possible values:
          - cls:        Select the CLS token as embedding
          - mean:       Apply Mean pooling to the model embeddings
          - splade:     Apply SPLADE (Sparse Lexical and Expansion) to the model embeddings. This option is only available if the loaded model is a `ForMaskedLM` Transformer model
          - last-token: Select the last token as embedding

      --max-concurrent-requests <MAX_CONCURRENT_REQUESTS>
          The maximum amount of concurrent requests for this particular deployment. Having a low limit will refuse clients requests instead of having them wait for too long and is usually good to handle backpressure correctly

          [env: MAX_CONCURRENT_REQUESTS=]
          [default: 512]

      --max-batch-tokens <MAX_BATCH_TOKENS>
          **IMPORTANT** This is one critical control to allow maximum usage of the available hardware.

          This represents the total amount of potential tokens within a batch.

          For `max_batch_tokens=1000`, you could fit `10` queries of `total_tokens=100` or a single query of `1000` tokens.

          Overall this number should be the largest possible until the model is compute bound. Since the actual memory overhead depends on the model implementation, text-embeddings-inference cannot infer this number automatically.

          [env: MAX_BATCH_TOKENS=]
          [default: 16384]

      --max-batch-requests <MAX_BATCH_REQUESTS>
          Optionally control the maximum number of individual requests in a batch

          [env: MAX_BATCH_REQUESTS=]

      --max-client-batch-size <MAX_CLIENT_BATCH_SIZE>
          Control the maximum number of inputs that a client can send in a single request

          [env: MAX_CLIENT_BATCH_SIZE=]
          [default: 32]

      --auto-truncate
          Control automatic truncation of inputs that exceed the model's maximum supported size. Defaults to `true` (truncation enabled). Set to `false` to disable truncation; when disabled and the model's maximum input length exceeds `--max-batch-tokens`, the server will refuse to start with an error instead of silently truncating sequences.

          Unused for gRPC servers

          [env: AUTO_TRUNCATE=]

      --default-prompt-name <DEFAULT_PROMPT_NAME>
          The name of the prompt that should be used by default for encoding. If not set, no prompt will be applied.

          Must be a key in the `sentence-transformers` configuration `prompts` dictionary.

          For example if ``default_prompt_name`` is "query" and the ``prompts`` is {"query": "query: ", ...}, then the sentence "What is the capital of France?" will be encoded as "query: What is the capital of France?" because the prompt text will be prepended before any text to encode.

          The argument '--default-prompt-name <DEFAULT_PROMPT_NAME>' cannot be used with '--default-prompt <DEFAULT_PROMPT>`

          [env: DEFAULT_PROMPT_NAME=]

      --default-prompt <DEFAULT_PROMPT>
          The prompt that should be used by default for encoding. If not set, no prompt will be applied.

          For example if ``default_prompt`` is "query: " then the sentence "What is the capital of France?" will be encoded as "query: What is the capital of France?" because the prompt text will be prepended before any text to encode.

          The argument '--default-prompt <DEFAULT_PROMPT>' cannot be used with '--default-prompt-name <DEFAULT_PROMPT_NAME>`

          [env: DEFAULT_PROMPT=]

      --dense-path <DENSE_PATH>
          Optionally, define the path to the Dense module required for some embedding models.

          Some embedding models require an extra `Dense` module which contains a single Linear layer and an activation function. By default, those `Dense` modules are stored under the `2_Dense` directory, but there might be cases where different `Dense` modules are provided, to convert the pooled embeddings into different dimensions, available as `2_Dense_<dims>` e.g. https://huggingface.co/NovaSearch/stella_en_400M_v5.

          Note that this argument is optional, only required to be set if there is no `modules.json` file or when you want to override a single Dense module path, only when running with the `candle` backend.

          [env: DENSE_PATH=]

      --hf-token <HF_TOKEN>
          Your Hugging Face Hub token. If neither `--hf-token` nor `HF_TOKEN` are set, the token will be read from the `$HF_HOME/token` path, if it exists. This ensures access to private or gated models, and allows for a more permissive rate limiting

          [env: HF_TOKEN=]

      --hostname <HOSTNAME>
          The IP address to listen on

          [env: HOSTNAME=]
          [default: 0.0.0.0]

      -p, --port <PORT>
          The port to listen on

          [env: PORT=]
          [default: 3000]

      --uds-path <UDS_PATH>
          The name of the unix socket some text-embeddings-inference backends will use as they communicate internally with gRPC

          [env: UDS_PATH=]
          [default: /tmp/text-embeddings-inference-server]

      --huggingface-hub-cache <HUGGINGFACE_HUB_CACHE>
          The location of the huggingface hub cache. Used to override the location if you want to provide a mounted disk for instance

          [env: HUGGINGFACE_HUB_CACHE=]

      --payload-limit <PAYLOAD_LIMIT>
          Payload size limit in bytes

          Default is 2MB

          [env: PAYLOAD_LIMIT=]
          [default: 2000000]

      --api-key <API_KEY>
          Set an api key for request authorization.

          By default the server responds to every request. With an api key set, the requests must have the Authorization header set with the api key as Bearer token.

          [env: API_KEY=]

      --json-output
          Outputs the logs in JSON format (useful for telemetry)

          [env: JSON_OUTPUT=]

      --disable-spans
          Whether or not to include the log trace through spans

          [env: DISABLE_SPANS=]

      --otlp-endpoint <OTLP_ENDPOINT>
          The grpc endpoint for opentelemetry. Telemetry is sent to this endpoint as OTLP over gRPC. e.g. `http://localhost:4317`

          [env: OTLP_ENDPOINT=]

      --otlp-service-name <OTLP_SERVICE_NAME>
          The service name for opentelemetry. e.g. `text-embeddings-inference.server`

          [env: OTLP_SERVICE_NAME=]
          [default: text-embeddings-inference.server]

      --prometheus-port <PROMETHEUS_PORT>
          The Prometheus port to listen on

          [env: PROMETHEUS_PORT=]
          [default: 9000]

      --cors-allow-origin <CORS_ALLOW_ORIGIN>
          Unused for gRPC servers

          [env: CORS_ALLOW_ORIGIN=]

  -h, --help
          Print help (see a summary with '-h')

  -V, --version
          Print version

Docker Images

Text Embeddings Inference ships with multiple Docker images that you can use to target a specific backend:
ArchitecturePlatformImage
CPUx8664ghcr.io/huggingface/text-embeddings-inference:cpu-1.9
CPUaarch64ghcr.io/huggingface/text-embeddings-inference:cpu-arm64-1.9
Voltax8664NOT SUPPORTED
Turing (T4, RTX 2000 series, ...)x8664ghcr.io/huggingface/text-embeddings-inference:turing-1.9 (experimental)
Ampere 8.0 (A100, A30)x8664ghcr.io/huggingface/text-embeddings-inference:1.9
Ampere 8.6 (A10, A40, ...)x8664ghcr.io/huggingface/text-embeddings-inference:86-1.9
Ada Lovelace (RTX 4000 series, ...)x8664ghcr.io/huggingface/text-embeddings-inference:89-1.9
Hopper (H100)x8664ghcr.io/huggingface/text-embeddings-inference:hopper-1.9
Blackwell 10.0 (B200, GB200, ...)x8664ghcr.io/huggingface/text-embeddings-inference:100-1.9 (experimental)
Blackwell 12.0 (GeForce RTX 50X0, ...)x8664ghcr.io/huggingface/text-embeddings-inference:120-1.9 (experimental)
Blackwell 12.1 (DGX Spark GB10, ...)multighcr.io/huggingface/text-embeddings-inference:121-1.9 (experimental)

Warning: Flash Attention is turned off by default for the Turing image as it suffers from precision issues. You can turn Flash Attention v1 ON by using the USE_FLASH_ATTENTION=True environment variable.

API documentation

You can consult the OpenAPI documentation of the text-embeddings-inference REST API using the /docs route. The Swagger UI is also available at: https://huggingface.github.io/text-embeddings-inference.

Using a private or gated model

You have the option to utilize the HF_TOKEN environment variable for configuring the token employed by text-embeddings-inference. This allows you to gain access to protected resources.
For example:
  1. Go to https://huggingface.co/settings/tokens
  2. Copy your CLI READ token
  3. Export HF_TOKEN=<your CLI READ token>

or with Docker:
model=<your private model>
volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run
token=<your CLI READ token>

docker run --gpus all -e HF_TOKEN=$token -p 8080:80 -v $volume:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 --model-id $model

Air gapped deployment

To deploy Text Embeddings Inference in an air-gapped environment, first download the weights and then mount them inside the container using a volume.
For example:
# (Optional) create a `models` directory
mkdir models
cd models

# Make sure you have git-lfs installed (https://git-lfs.com)
git lfs install
git clone https://huggingface.co/Qwen/Qwen3-Embedding-0.6B

# Set the models directory as the volume path
volume=$PWD

# Mount the models directory inside the container with a volume and set the model ID
docker run --gpus all -p 8080:80 -v $volume:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 --model-id /data/Qwen3-Embedding-0.6B

Using Re-rankers models

text-embeddings-inference v0.4.0 added support for CamemBERT, RoBERTa, XLM-RoBERTa, and GTE Sequence Classification models. Re-rankers models are Sequence Classification cross-encoders models with a single class that scores the similarity between a query and a text.
See this blogpost by the LlamaIndex team to understand how you can use re-rankers models in your RAG pipeline to improve downstream performance.
model=BAAI/bge-reranker-large
volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run

docker run --gpus all -p 8080:80 -v $volume:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 --model-id $model

And then you can rank the similarity between a query and a list of texts with:
curl 127.0.0.1:8080/rerank \
    -X POST \
    -d '{"query": "What is Deep Learning?", "texts": ["Deep Learning is not...", "Deep learning is..."]}' \
    -H 'Content-Type: application/json'

Using Sequence Classification models

You can also use classic Sequence Classification models like SamLowe/roberta-base-go_emotions:
model=SamLowe/roberta-base-go_emotions
volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run

docker run --gpus all -p 8080:80 -v $volume:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 --model-id $model

Once you have deployed the model you can use the predict endpoint to get the emotions most associated with an input:
curl 127.0.0.1:8080/predict \
    -X POST \
    -d '{"inputs":"I like you."}' \
    -H 'Content-Type: application/json'

Using SPLADE pooling

You can choose to activate SPLADE pooling for Bert and Distilbert MaskedLM architectures:
model=naver/efficient-splade-VI-BT-large-query
volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run

docker run --gpus all -p 8080:80 -v $volume:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 --model-id $model --pooling splade

Once you have deployed the model you can use the /embed_sparse endpoint to get the sparse embedding:
curl 127.0.0.1:8080/embed_sparse \
    -X POST \
    -d '{"inputs":"I like you."}' \
    -H 'Content-Type: application/json'

Distributed Tracing

text-embeddings-inference is instrumented with distributed tracing using OpenTelemetry. You can use this feature by setting the address to an OTLP collector with the --otlp-endpoint argument.

gRPC

text-embeddings-inference offers a gRPC API as an alternative to the default HTTP API for high performance deployments. The API protobuf definition can be found here.
You can use the gRPC API by adding the -grpc tag to any TEI Docker image. For example:
model=Qwen/Qwen3-Embedding-0.6B
volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run

docker run --gpus all -p 8080:80 -v $volume:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cuda-1.9-grpc --model-id $model

grpcurl -d '{"inputs": "What is Deep Learning"}' -plaintext 0.0.0.0:8080 tei.v1.Embed/Embed

Local install

Apple Silicon (Homebrew)

On Apple Silicon (M1/M2/M3/M4), you can install a prebuilt binary via Homebrew:
brew install text-embeddings-inference

Then launch Text Embeddings Inference with Metal acceleration:
model=Qwen/Qwen3-Embedding-0.6B

text-embeddings-router --model-id $model --port 8080

CPU

You can also opt to install text-embeddings-inference locally.
First install Rust:
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh

Then run:
# On x86 with ONNX backend (recommended)
cargo install --path router -F ort
# On x86 with Intel backend
cargo install --path router -F mkl
# On M1 or M2
cargo install --path router -F metal

You can now launch Text Embeddings Inference on CPU with:
model=Qwen/Qwen3-Embedding-0.6B

text-embeddings-router --model-id $model --port 8080

Note: on some machines, you may also need the OpenSSL libraries and gcc. On Linux machines, run:
sudo apt-get install libssl-dev gcc -y

CUDA

GPUs with CUDA compute capabilities < 7.5 are not supported (V100, Titan V, GTX 1000 series, ...).
Make sure you have CUDA and the NVIDIA drivers installed. NVIDIA drivers on your device need to be compatible with CUDA version 12.2 or higher. You also need to add the NVIDIA binaries to your path:
export PATH=$PATH:/usr/local/cuda/bin

Then run the following (might take a while as it needs to compile the CUDA kernels):
# On Turing GPUs (T4, RTX 2000 series ... )
cargo install --path router -F candle-cuda-turing

# On Ampere, Ada Lovelace, Hopper and Blackwell
cargo install --path router -F candle-cuda

You can now launch Text Embeddings Inference on GPU as follows:
model=Qwen/Qwen3-Embedding-0.6B

text-embeddings-router --model-id $model --port 8080

Docker

You can build the CPU container with Docker as:
docker build -f Dockerfile .

To build the CUDA containers, you need to know the compute cap of the GPU you will be using at runtime, to build the image accordingly:
# Get submodule dependencies
git submodule update --init

# Example for Turing (T4, RTX 2000 series, ...)
runtime_compute_cap=75

# Example for Ampere (A100, ...)
runtime_compute_cap=80

# Example for Ampere (A10, ...)
runtime_compute_cap=86

# Example for Ada Lovelace (RTX 4000 series, ...)
runtime_compute_cap=89

# Example for Hopper (H100, ...)
runtime_compute_cap=90

# Example for Blackwell (B200, GB200, ...)
runtime_compute_cap=100

# Example for Blackwell (GeForce RTX 50X0, RTX PRO 6000, ...)
runtime_compute_cap=120

# Example for Blackwell GB10 (DGX Spark)
runtime_compute_cap=121

docker build . -f Dockerfile-cuda --build-arg CUDA_COMPUTE_CAP=$runtime_compute_cap

ARM64 / aarch64

CPU-only (Apple Silicon, Ampere, Graviton)

For ARM64 hosts without NVIDIA GPUs, use the CPU Dockerfile. Inference runs on CPU cores only (no Metal/MPS support via Docker).
docker build . -f Dockerfile-arm64 --platform=linux/arm64

CUDA on ARM64 (DGX Spark, Jetson)

For ARM64 hosts with NVIDIA GPUs, build Dockerfile-cuda with the appropriate compute capability and --platform linux/arm64:
# DGX Spark (GB10, sm_121)
docker build . -f Dockerfile-cuda \
  --build-arg CUDA_COMPUTE_CAP=121 \
  --platform linux/arm64

# Future ARM64 + Blackwell devices (sm_120)
docker build . -f Dockerfile-cuda \
  --build-arg CUDA_COMPUTE_CAP=120 \
  --platform linux/arm64

AMD Instinct GPUs (ROCm)

TEI supports AMD Instinct GPUs (MI200, MI300 series) via ROCm.
model=BAAI/bge-base-en-v1.5
volume=$PWD/data

docker run \
  --device /dev/kfd --device /dev/dri \
  --group-add video \
  --ipc=host \
  -p 8080:80 \
  -v $volume:/data \
  --pull always \
  ghcr.io/huggingface/text-embeddings-inference:rocm-latest \
  --model-id $model --dtype bfloat16

For full setup instructions, see the AMD Instinct GPU guide.

Examples

Serve TEI Embeddings on your own domain behind Caddy, Nginx or Traefik. Fill in your domain and copy the result. It's a starting point, some apps need their own base URL or extra headers set too.

Proxying tei-embeddings.example.com to http://tei:5000

Add this to your Caddyfile

tei-embeddings.example.com {
	reverse_proxy http://tei:5000
}

Check the logs first

Nine times out of ten the logs tell you exactly what went wrong.

  • In Portainer, go to Containers, click the container, then Logs. Or run docker logs <container>
  • Exit codes help too: 137 means killed, usually out of memory. 126 or 127 means the command inside the image is broken.

Port already in use

If deployment fails with "Bind for 0.0.0.0:5000 failed: port is already allocated", something else on your server is using that port.

  • Find what's using it: sudo ss -tlnp | grep :5000
  • Stop the other service, or pick a different host port. In 5000:5000 only the left number is yours to change, the right one belongs to the app.

Running but the page won't load

The container is up but nothing appears in your browser.

  • Use your server's real IP: http://your-server-ip:5000. The 0.0.0.0 link Portainer shows isn't a real address.
  • Give it a minute after first deploy, tei can take a while to initialise.
  • Make sure your firewall allows the port, e.g. sudo ufw allow 5000

Image won't pull

Test the pull directly on the host: docker pull ghcr.io/huggingface/text-embeddings-inference:cpu-1.9

  • "manifest unknown" means the tag no longer exists.
  • "toomanyrequests" is the Docker Hub rate limit. Log in with docker login to raise it.
  • "no space left on device" means a full disk. Reclaim space with docker system prune

"exec format error"

This means the image was built for a different CPU architecture than your server.

  • This image supports: amd64
  • Check yours with uname -m: x86_64 is amd64, aarch64 is arm64. Raspberry Pi and other ARM boards are the usual culprits.

Container keeps restarting

The unless-stopped restart policy relaunches the app after every crash, so the real error can scroll past.

  • Check the logs right after a restart, the last few lines before it died are the useful ones.
  • Get the exit code with docker inspect <container> --format '{{.State.ExitCode}}'
  • Still stuck? Redeploy once with the restart policy set to no so the failure stays visible.

Stack won't deploy

Compose stacks fail fast on small mistakes, and Portainer shows the reason just above the editor.

  • YAML only accepts spaces for indentation, a single tab breaks the whole file.

Raise an issue

Found something which isn't working as it should? Here's how to report it.

A Compose stack

TEI Embeddings is a Compose stack, a set of containers defined in one file and brought up together by Portainer, then started and stopped as a single app.

The app image

An image is the app packed up ready to go, everything TEI Embeddings needs bundled into one download. This template pulls ghcr.io/huggingface/text-embeddings-inference:cpu-1.9, which Docker fetches once (about 228 MB) and then starts your own copy from.

Where the image comes from

Docker pulls its images from registries, public libraries of ready-built apps. TEI Embeddings's comes from the GitHub Container Registry, published by huggingface.

Version tags

The bit after the colon in the image name is the version tag. This one pins cpu-1.9, so every redeploy gives you that exact build until you bump it yourself.

Which machines it runs on

Every image is built for particular CPU types. This one ships for amd64, so it runs on regular x86 PCs and servers, though not ARM boards like a Raspberry Pi.

Ports

A port is the door the app answers on. A mapping like 5000:5000 means it's reachable on port 5000 of your server, where the left number is yours to change and the right one belongs to the app. It opens:

  • 5000:5000

Volumes

A volume is where TEI Embeddings keeps its files so they survive an update or a restart. Without one, anything it saves would sit inside the container and vanish the moment it's recreated. This template mounts:

  • /data kept in the modelcache volume Docker manages

Environment variables

Environment variables are the settings you hand over when you deploy, things like a password or a timezone. TEI Embeddings takes 3 of them, all with defaults you can leave alone or tweak:

  • PORT, defaults to 5000
  • MODEL_ID, defaults to BAAI/bge-small-en-v1.5
  • HUGGINGFACE_HUB_CACHE, defaults to /data

Restart policy

The restart policy here is unless-stopped, so Docker restarts TEI Embeddings after a crash or reboot, but leaves it off when you stop it on purpose. You can change this on the deploy screen. The choices are no (never restart), on-failure (only after a crash), unless-stopped (restart unless you stop it), and always (bring it back no matter what).

Networking

Nothing custom is set, so TEI Embeddings sits on Docker's default bridge network: its own private space that reaches the outside world only through the ports it publishes.

Container name

Once it's deployed, Portainer names the container tei. That's what you'll spot in the containers list and use in commands like docker logs tei.

Platform

The platform is linux, the kind of system the container is built to run on. Docker and Portainer handle this on a normal Linux server.

Open source license

TEI Embeddings is open source, released under the Apache-2.0 license. In plain terms the code is out in the open, so you're free to run it and change it to fit what you need.

Portainer app templates

Zooming out, this whole page comes from a Portainer app template: a short recipe telling Portainer how to set TEI Embeddings up. Add the template list to Portainer once, then deploying TEI Embeddings is a click rather than a wall of config.