Add support for idefics3 (SmolVLM) #1059

xenova · 2024-11-29T03:01:15Z

Example usage:

import {
  AutoProcessor,
  AutoModelForVision2Seq,
  load_image,
} from "@huggingface/transformers";

// Initialize processor and model
const model_id = "HuggingFaceTB/SmolVLM-Instruct";
const processor = await AutoProcessor.from_pretrained(model_id);
const model = await AutoModelForVision2Seq.from_pretrained(model_id, {
  dtype: {
    embed_tokens: "fp16", // "fp32", "fp16", "q8"
    vision_encoder: "q4", // "fp32", "fp16", "q8", "q4", "q4f16"
    decoder_model_merged: "q4", // "q8", "q4", "q4f16"
  }
});

// Load images
const image1 = await load_image("https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg");
const image2 = await load_image("https://huggingface.co/spaces/merve/chameleon-7b/resolve/main/bee.jpg");

// Create input messages
const messages = [
  {
    role: "user",
    content: [
      { type: "image" },
      { type: "image" },
      { type: "text", text: "Can you describe the two images?" },
    ],
  },
];

// Prepare inputs
const text = processor.apply_chat_template(messages, { add_generation_prompt: true });
const inputs = await processor(text, [image1, image2], {
  // Set `do_image_splitting: true` to split images into multiple patches.
  // NOTE: This uses more memory, but can provide more accurate results.
  do_image_splitting: false,
});

// Generate outputs
const generated_ids = await model.generate({
  ...inputs,
  max_new_tokens: 500,
});
const generated_texts = processor.batch_decode(
  generated_ids.slice(null, [inputs.input_ids.dims.at(-1), null]),
  { skip_special_tokens: true },
);
console.log(generated_texts[0]);
// ' In the first image, there is a green statue of liberty on a pedestal in the middle of the water. The water is surrounded by trees and buildings in the background. In the second image, there are pink and red flowers with a bee on the pink flower.'

HuggingFaceDocBuilderDev · 2024-11-29T03:03:44Z

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

…lacement

xenova added 2 commits November 29, 2024 02:59

[WIP] Add support for idefics3 (SmolVLM)

3a02a52

Cleanup

293e378

xenova marked this pull request as draft November 29, 2024 03:01

xenova added 16 commits November 29, 2024 11:34

Update DataTypeMap with 4-bit data types

ffa43f2

Format the model inputs before logging to console

01dd2d1

Use QUInt8 when quantizing models produced by onnxruntime-genai

f95475f

auto dtype selection

3aab729

Export load_image helper function

33cbbd7

Add listed support for Idefics3

832a7ce

Add support for batched 2d images in idefics3 processor

f9c59f4

Update unit tests

2e0c5bd

Add another unit test to ensure correctness of pixel attention mask p…

cadf689

…lacement

Move image tokens out of call function

d3205de

Formatting

2fd4838

Improve auto selection logic

84c5b7c

Return correct pixel_attention_mask

45de96a

Update pixel_attention_mask unit test

f88500c

Formatting

b448b4e

Add idefics3 unit tests

14a0968

xenova changed the title ~~[WIP] Add support for idefics3 (SmolVLM)~~ Add support for idefics3 (SmolVLM) Dec 2, 2024

xenova marked this pull request as ready for review December 2, 2024 20:56

Increase idefics processor unit test timeout

e836628

xenova merged commit 11db949 into main Dec 2, 2024
4 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Add support for idefics3 (SmolVLM) #1059

Add support for idefics3 (SmolVLM) #1059

xenova commented Nov 29, 2024 •

edited

Loading

HuggingFaceDocBuilderDev commented Nov 29, 2024

Add support for idefics3 (SmolVLM) #1059

Add support for idefics3 (SmolVLM) #1059

Conversation

xenova commented Nov 29, 2024 • edited Loading

HuggingFaceDocBuilderDev commented Nov 29, 2024

xenova commented Nov 29, 2024 •

edited

Loading