Referenced from https://github.com/triton-inference-server/python_backend
- Run triton server container to start up the service and a client that is calling the service
# in seperate terminals
bash run_example_server.sh
bash run_example_client.sh
- Edit docker volume location of
run_simple_server.shandrun_simple_client.shaccordingly - Run triton server container to start up the service and a client that is calling the service in seperate terminals
bash run_simple_server.sh
bash run_simple_client.sh
The model used here is the resnet18 from https://pytorch.org/vision/stable/models.html. Weights are downloaded via setting the pretrained flag and can be found on your system in ~/.cache/torch/hub/checkpoints/
If you are using a python version different from python 3.8, you will have to refer to point 1 "Building Custom Python Backend Stub" of this link.
- Take note that you will have to have a gcc version <= 8. You can refer to this link
- You will also need nvcc so you can get it by installing cudatoolkit on the host. Just
sudo apt-get install cudatoolkit
If you are using python 3.8 but you have additional libraries required, you will have to refer to point 2 "Packaging the Conda Environment" of this link.
Source files:
pb_utils
http client
Triton configurations:
.pbtxt
http protocol
figure out how to send image via http form data