[Feature request] Setting model size and max concurrency specifically for each model (Triton). #93

haiminh2001 · 2024-09-14T14:28:56Z

message LoadModelResponse {
    // OPTIONAL - If nontrivial cost is involved in
    // determining the size, return 0 here and
    // do the sizing in the modelSize function
    uint64 sizeInBytes = 1;

    // EXPERIMENTAL - Applies only if limitModelConcurrency = true
    // was returned from runtimeStatus rpc.
    // See RuntimeStatusResponse.limitModelConcurrency for more detail
    uint32 maxConcurrency = 2;
}

Hi, in the model-runtime.proto, the LoadModelResponse specify the model size in bytes and the max concurrency of the model. Currently, the size in bytes is hard-coded as the size of model files, which may be reasonable for Deep Learning weights but inaccurate for example, triton python backend. In addition, each model should indeed have different max concurrency.
Therefore, I propose that the adapter perhaps can read these configurations from a separate config file within the model folder (just like the config.pbtxt file) to override these configurations.
I am open to create a PR.

The text was updated successfully, but these errors were encountered:

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Feature request] Setting model size and max concurrency specifically for each model (Triton). #93

[Feature request] Setting model size and max concurrency specifically for each model (Triton). #93

haiminh2001 commented Sep 14, 2024

[Feature request] Setting model size and max concurrency specifically for each model (Triton). #93

[Feature request] Setting model size and max concurrency specifically for each model (Triton). #93

Comments

haiminh2001 commented Sep 14, 2024