[
](ABOUT_AINYAN.md)
[](https://github.com/ailia-ai/ailia-models/stargazers)
[](https://pypi.org/project/ailia/)


The collection of pre-trained, state-of-the-art AI models: 418 models covering object detection, speech recognition, image generation, LLMs and more — all runnable from the same simple CLI.
[Tutorial](TUTORIAL.md) · [チュートリアル](TUTORIAL_jp.md) · [Google Colaboratory](https://www.ailia.ai/launch_to_colab) · [Documentation](https://docs.ailia.ai/en/) · [deepwiki](https://deepwiki.com/ailia-ai/ailia-models) · [Update history](https://github.com/ailia-ai/ailia-models/wiki)
Every model works the same way: no arguments needed, weights download automatically.
```
pip3 install ailia
git clone https://github.com/ailia-ai/ailia-models
cd ailia-models
pip3 install -r requirements.txt
cd object_detection/yolox
python3 yolox.py
```
# Models
419 models are available. Use [🔍 Search models](https://github.com/ailia-ai/ailia-models/find/master) to find a model by name.
| | Category | Model list |
|:---|:---|:---|
| [
](/action_recognition/) | [Action recognition](/action_recognition/) | [va-cnn](/action_recognition/va-cnn/), [st-gcn](/action_recognition/st_gcn/), [mars](/action_recognition/mars/), [ax_action_recognition](/action_recognition/ax_action_recognition/), [driver-action-recognition-adas](/action_recognition/driver-action-recognition-adas/), [action_clip](/action_recognition/action_clip/) |
| [
](/anomaly_detection/) | [Anomaly detection](/anomaly_detection/) | [mahalanobisad](/anomaly_detection/mahalanobisad/), [spade-pytorch](/anomaly_detection/spade-pytorch/), [padim](/anomaly_detection/padim/), [patchcore](/anomaly_detection/patchcore/), [glass](/anomaly_detection/glass/) |
| | [Audio language model](/audio_language_model/) | [qwen_audio](/audio_language_model/qwen_audio/) |
| | [Audio processing](/audio_processing/) | **Audio classification**: [crnn_audio_classification](/audio_processing/crnn_audio_classification/), [audioset_tagging_cnn](/audio_processing/audioset_tagging_cnn/), [transformer-cnn-emotion-recognition](/audio_processing/transformer-cnn-emotion-recognition/), [microsoft clap](/audio_processing/msclap/), [clap](/audio_processing/clap/)
**Music enhancement**: [hifigan](/audio_processing/hifigan/), [deep music enhancer](/audio_processing/deep-music-enhancer/)
**Music generation**: [pytorch_wavenet](/audio_processing/pytorch_wavenet/)
**Noise reduction**: [rnnoise](/audio_processing/rnnoise/), [voicefilter](/audio_processing/voicefilter/), [unet_source_separation](/audio_processing/unet_source_separation/), [demucs](/audio_processing/demucs/), [dtln](/audio_processing/dtln/), [voicesplit](/audio_processing/voicesplit/), [audiosep](/audio_processing/audiosep/)
**Phoneme alignment**: [narabas](/audio_processing/narabas/)
**Pitch detection**: [crepe](/audio_processing/crepe/)
**Speaker diarization**: [pyannote-audio](/audio_processing/pyannote-audio/), [auto_speech](/audio_processing/auto_speech/), [wespeaker](/audio_processing/wespeaker/)
**Speech to text**: [deepspeech2](/audio_processing/deepspeech2/), [whisper](/audio_processing/whisper/), [reazon_speech](/audio_processing/reazon_speech/), [distil-whisper](/audio_processing/distil-whisper/), [sensevoice](/audio_processing/sensevoice/), [reazon_speech2](/audio_processing/reazon_speech2/), [kotoba-whisper](/audio_processing/kotoba-whisper/), [lite-whisper](/audio_processing/lite-whisper/)
**Text to speech**: [pytorch-dc-tts](/audio_processing/pytorch-dc-tts/), [tacotron2](/audio_processing/tacotron2/), [vall-e-x](/audio_processing/vall-e-x/), [Bert-VITS2](/audio_processing/bert-vits2/), [gpt-sovits](/audio_processing/gpt-sovits/), [gpt-sovits-v2](/audio_processing/gpt-sovits-v2/), [cosyvoice2](/audio_processing/cosyvoice2/), [gpt-sovits-v3](/audio_processing/gpt-sovits-v3/), [gpt-sovits-v2-pro](/audio_processing/gpt-sovits-v2-pro/), [qwen3-tts](/audio_processing/qwen3-tts/)
**Voice activity detection**: [silero-vad](/audio_processing/silero-vad/)
**Voice conversion**: [rvc](/audio_processing/rvc/) |
| [
](/autonomous_driving/) | [Autonomous driving](/autonomous_driving/) | [bevformer](/autonomous_driving/bevformer/), [segformer](/autonomous_driving/segformer/), [uniad](/autonomous_driving/uniad/) |
| [
](/background_removal/) | [Background removal](/background_removal/) | [deep-image-matting](/background_removal/deep-image-matting/), [indexnet](/background_removal/indexnet/), [U-2-Net](/background_removal/u2net/), [u2net-portrait-matting](/background_removal/u2net-portrait-matting/), [u2net-human-seg](/background_removal/u2net-human-seg/), [cascade_psp](/background_removal/cascade_psp/), [rembg](/background_removal/rembg/), [gfm](/background_removal/gfm/), [modnet](/background_removal/modnet/), [background_matting_v2](/background_removal/background_matting_v2/), [dis_seg](/background_removal/dis_seg/) |
| [
](/crowd_counting/) | [Crowd counting](/crowd_counting/) | [crowdcount-cascaded-mtl](/crowd_counting/crowdcount-cascaded-mtl/), [c-3-framework](/crowd_counting/c-3-framework/) |
| [
](/deep_fashion/) | [Deep fashion](/deep_fashion/) | [fashionai-key-points-detection](/deep_fashion/fashionai-key-points-detection/), [person-attributes-recognition-crossroad](/deep_fashion/person-attributes-recognition-crossroad/), [clothing-detection](/deep_fashion/clothing-detection/), [mmfashion](/deep_fashion/mmfashion/), [mmfashion_tryon](/deep_fashion/mmfashion_tryon/), [mmfashion_retrieval](/deep_fashion/mmfashion_retrieval/) |
| [
](/depth_estimation/) | [Depth estimation](/depth_estimation/) | [fcrn-depthprediction](/depth_estimation/fcrn-depthprediction/), [monodepth2](/depth_estimation/monodepth2/), [fast-depth](/depth_estimation/fast-depth/), [midas](/depth_estimation/midas/), [hitnet](/depth_estimation/hitnet/), [lap-depth](/depth_estimation/lap-depth/), [mobilestereonet](/depth_estimation/mobilestereonet/), [crestereo](/depth_estimation/crestereo/), [zoe_depth](/depth_estimation/zoe_depth/), [depth_anything](/depth_estimation/depth_anything/), [depth_anything_v2](/depth_estimation/depth_anything_v2/), [depth_pro](/depth_estimation/depth_pro/), [depth_anything_v3](/depth_estimation/depth_anything_v3/) |
| [
](/diffusion/) | [Diffusion](/diffusion/) | **Text to image**: [latent-diffusion-txt2img](/diffusion/latent-diffusion-txt2img/), [stable-diffusion-txt2img](/diffusion/stable-diffusion-txt2img/), [anything_v3](/diffusion/anything_v3/), [control_net](/diffusion/control_net/), [sdxl](/diffusion/sdxl/), [latent-consistency-models](/diffusion/latent-consistency-models/), [sd-turbo](/diffusion/sd-turbo/), [sdxl-turbo](/diffusion/sdxl-turbo/), [depth_anything_controlnet](/diffusion/depth_anything_controlnet/), [latentsync](/diffusion/latentsync/)
**Text to audio**: [riffusion](/diffusion/riffusion/)
**Others**: [latent-diffusion-inpainting](/diffusion/latent-diffusion-inpainting/), [latent-diffusion-superresolution](/diffusion/latent-diffusion-superresolution/), [DA-CLIP](/diffusion/daclip-sde/), [marigold](/diffusion/marigold/) |
| [
](/face_detection/) | [Face detection](/face_detection/) | [mtcnn](/face_detection/mtcnn/), [yolov1-face](/face_detection/yolov1-face/), [face-detection-adas](/face_detection/face-detection-adas/), [retinaface](/face_detection/retinaface/), [blazeface](/face_detection/blazeface/), [yolov3-face](/face_detection/yolov3-face/), [face-mask-detection](/face_detection/face-mask-detection/), [dbface](/face_detection/dbface/), [anime-face-detector](/face_detection/anime-face-detector/) |
| [
](/face_identification/) | [Face identification](/face_identification/) | [facenet_pytorch](/face_identification/facenet_pytorch/), [insightface](/face_identification/insightface/), [vggface2](/face_identification/vggface2/), [arcface](/face_identification/arcface/), [cosface](/face_identification/cosface/) |
| [
](/face_recognition/) | [Face recognition](/face_recognition/) | **Age gender estimation**: [face_classification](/face_recognition/face_classification/), [age-gender-recognition-retail](/face_recognition/age-gender-recognition-retail/), [mivolo](/face_recognition/mivolo/)
**Emotion recognition**: [ferplus](/face_recognition/ferplus/), [hsemotion](/face_recognition/hsemotion/)
**Gaze estimation**: [gazeml](/face_recognition/gazeml/), [mediapipe_iris](/face_recognition/mediapipe_iris/), [gazelle](/face_recognition/gazelle/), [ax_gaze_estimation](/face_recognition/ax_gaze_estimation/)
**Head pose estimation**: [hopenet](/face_recognition/hopenet/), [6d_repnet](/face_recognition/6d_repnet/), [L2CS_Net](/face_recognition/l2cs_net/), [6d_repnet_360](/face_recognition/6d_repnet_360/)
**Keypoint detection**: [face_alignment](/face_recognition/face_alignment/), [prnet](/face_recognition/prnet/), [facemesh](/face_recognition/facemesh/), [facial_feature](/face_recognition/facial_feature/), [3ddfa](/face_recognition/3ddfa/), [facemesh_v2](/face_recognition/facemesh_v2/)
**Others**: [face-anti-spoofing](/face_recognition/face-anti-spoofing/), [ax_facial_features](/face_recognition/ax_facial_features/) |
| [
](/face_restoration/) | [Face restoration](/face_restoration/) | [gfpgan](/face_restoration/gfpgan/), [codeformer](/face_restoration/codeformer/) |
| [
](/face_swapping/) | [Face swapping](/face_swapping/) | [deepfacelive](/face_swapping/deepfacelive/), [sber-swap](/face_swapping/sber-swap/), [facefusion](/face_swapping/facefusion/) |
| [
](/feature_extraction/) | [Feature extraction](/feature_extraction/) | [dinov3](/feature_extraction/dinov3/) |
| [
](/frame_interpolation/) | [Frame interpolation](/frame_interpolation/) | [cain](/frame_interpolation/cain/), [rife](/frame_interpolation/rife/), [flavr](/frame_interpolation/flavr/), [film](/frame_interpolation/film/) |
| [
](/generative_adversarial_networks/) | [Generative adversarial networks](/generative_adversarial_networks/) | [pytorch-gan](/generative_adversarial_networks/pytorch-gan/), [lipgan](/generative_adversarial_networks/lipgan/), [council-gan](/generative_adversarial_networks/council-gan/), [sam](/generative_adversarial_networks/sam/), [encoder4editing](/generative_adversarial_networks/encoder4editing/), [restyle-encoder](/generative_adversarial_networks/restyle-encoder/), [SadTalker](/generative_adversarial_networks/sadtalker/), [live_portrait](/generative_adversarial_networks/live_portrait/) |
| [
](/hand_detection/) | [Hand detection](/hand_detection/) | [hand_detection_pytorch](/hand_detection/hand_detection_pytorch/), [yolov3-hand](/hand_detection/yolov3-hand/), [blazepalm](/hand_detection/blazepalm/) |
| [
](/hand_recognition/) | [Hand recognition](/hand_recognition/) | [hand3d](/hand_recognition/hand3d/), [v2v-posenet](/hand_recognition/v2v-posenet/), [minimal-hand](/hand_recognition/minimal-hand/), [blazehand](/hand_recognition/blazehand/), [hands_segmentation_pytorch](/hand_recognition/hands_segmentation_pytorch/) |
| [
](/image_captioning/) | [Image captioning](/image_captioning/) | [illustration2vec](/image_captioning/illustration2vec/), [image_captioning_pytorch](/image_captioning/image_captioning_pytorch/), [blip2](/image_captioning/blip2/) |
| [
](/image_classification/) | [Image classification](/image_classification/) | **CNN**: [alexnet](/image_classification/alexnet/), [vgg16](/image_classification/vgg16/), [googlenet](/image_classification/googlenet/), [resnet18](/image_classification/resnet18/), [resnet50](/image_classification/resnet50/), [inceptionv3](/image_classification/inceptionv3/), [inceptionv4](/image_classification/inceptionv4/), [wide_resnet50](/image_classification/wide_resnet50/), [mobilenetv2](/image_classification/mobilenetv2/), [mobilenetv3](/image_classification/mobilenetv3/), [efficientnet](/image_classification/efficientnet/), [efficientnetv2](/image_classification/efficientnetv2/), [imagenet21k](/image_classification/imagenet21k/), [mlp_mixer](/image_classification/mlp_mixer/), [volo](/image_classification/volo/), [convnext](/image_classification/convnext/), [mobileone](/image_classification/mobileone/)
**Transformer**: [vit](/image_classification/vit/), [clip](/image_classification/clip/), [swin-transformer](/image_classification/swin-transformer/), [japanese-clip](/image_classification/japanese-clip/), [japanese-stable-clip-vit-l-16](/image_classification/japanese-stable-clip-vit-l-16/), [siglip-multilingual](/image_classification/siglip-multilingual/), [clip-japanese-base](/image_classification/clip-japanese-base/), [siglip2](/image_classification/siglip2/)
**Specific task**: [weather-prediction-from-image](/image_classification/weather-prediction-from-image/), [partialconv](/image_classification/partialconv/) |
| [
](/image_inpainting/) | [Image inpainting](/image_inpainting/) | [inpainting-with-partial-conv](/image_inpainting/pytorch-inpainting-with-partial-conv/), [deepfillv2](/image_inpainting/deepfillv2/), [inpainting_gmcnn](/image_inpainting/inpainting_gmcnn/), [3d-photo-inpainting](/image_inpainting/3d-photo-inpainting/), [lama](/image_inpainting/lama/) |
| [
](/image_manipulation/) | [Image manipulation](/image_manipulation/) | [colorization](/image_manipulation/colorization/), [cnngeometric_pytorch](/image_manipulation/cnngeometric_pytorch/), [style2paints](/image_manipulation/style2paints/), [deblur_gan](/image_manipulation/deblur_gan/), [pytorch-superpoint](/image_manipulation/pytorch-superpoint/), [noise2noise](/image_manipulation/noise2noise/), [dfe](/image_manipulation/dfe/), [illnet](/image_manipulation/illnet/), [dewarpnet](/image_manipulation/dewarpnet/), [deep_white_balance](/image_manipulation/deep_white_balance/), [u2net_portrait](/image_manipulation/u2net_portrait/), [invertible_denoising_network](/image_manipulation/invertible_denoising_network/), [dfm](/image_manipulation/dfm/), [fbcnn](/image_manipulation/fbcnn/), [dehamer](/image_manipulation/dehamer/), [lightglue](/image_manipulation/lightglue/), [docshadow](/image_manipulation/docshadow/) |
| [
](/image_quality_assessment/) | [Image quality assessment](/image_quality_assessment/) | [aesthetic-predictor](/image_quality_assessment/aesthetic-predictor/) |
| [
](/image_restoration/) | [Image restoration](/image_restoration/) | [nafnet](/image_restoration/nafnet/) |
| [
](/image_segmentation/) | [Image segmentation](/image_segmentation/) | [pytorch-fcn](/image_segmentation/pytorch-fcn/), [pytorch-enet](/image_segmentation/pytorch-enet/), [tusimple-DUC](/image_segmentation/tusimple-DUC/), [pytorch-unet](/image_segmentation/pytorch-unet/), [deeplabv3](/image_segmentation/deeplabv3/), [pspnet-hair-segmentation](/image_segmentation/pspnet-hair-segmentation/), [swiftnet](/image_segmentation/swiftnet/), [hrnet_segmentation](/image_segmentation/hrnet_segmentation/), [hair_segmentation](/image_segmentation/hair_segmentation/), [paddleseg](/image_segmentation/paddleseg/), [human_part_segmentation](/image_segmentation/human_part_segmentation/), [semantic-segmentation-mobilenet-v3](/image_segmentation/semantic-segmentation-mobilenet-v3/), [suim](/image_segmentation/suim/), [yet-another-anime-segmenter](/image_segmentation/yet-another-anime-segmenter/), [dense_prediction_transformers](/image_segmentation/dense_prediction_transformers/), [group_vit](/image_segmentation/group_vit/), [pp_liteseg](/image_segmentation/pp_liteseg/), [anime-segmentation](/image_segmentation/anime-segmentation/), [yolov8-seg](/image_segmentation/yolov8-seg/), [segment-anything](/image_segmentation/segment-anything/), [grounded_sam](/image_segmentation/grounded_sam/), [fast_sam](/image_segmentation/fast_sam/), [mobile_sam](/image_segmentation/mobile_sam/), [edge_sam](/image_segmentation/edge_sam/), [segment-anything-2](/image_segmentation/segment-anything-2/), [yolov11-seg](/image_segmentation/yolov11-seg/), [segment-anything-3.1](/image_segmentation/segment-anything-3.1/) |
| [
](/landmark_classification/) | [Landmark classification](/landmark_classification/) | [places365](/landmark_classification/places365/), [landmarks_classifier_asia](/landmark_classification/landmarks_classifier_asia/) |
| [
](/line_segment_detection/) | [Line segment detection](/line_segment_detection/) | [dexined](/line_segment_detection/dexined/), [mlsd](/line_segment_detection/mlsd/) |
| [
](/low_light_image_enhancement/) | [Low light image enhancement](/low_light_image_enhancement/) | [agllnet](/low_light_image_enhancement/agllnet/), [drbn_skf](/low_light_image_enhancement/drbn_skf/) |
| | [Natural language processing](/natural_language_processing/) | **Bert**: [bert](/natural_language_processing/bert/), [bert_maskedlm](/natural_language_processing/bert_maskedlm/), [bert_question_answering](/natural_language_processing/bert_question_answering/)
**Embedding**: [sentence_transformers_japanese](/natural_language_processing/sentence_transformers_japanese/), [multilingual-e5](/natural_language_processing/multilingual-e5/), [glucose](/natural_language_processing/glucose/), [qwen3-embedding](/natural_language_processing/qwen3-embedding/), [ruri-v3](/natural_language_processing/ruri-v3/), [embeddinggemma](/natural_language_processing/embeddinggemma/)
**Error corrector**: [bert_insert_punctuation](/natural_language_processing/bert_insert_punctuation/), [bertjsc](/natural_language_processing/bertjsc/), [t5_whisper_medical](/natural_language_processing/t5_whisper_medical/)
**Grapheme to phoneme**: [g2p_en](/natural_language_processing/g2p_en/), [g2pw](/natural_language_processing/g2pw/), [soundchoice-g2p](/natural_language_processing/soundchoice-g2p/)
**Named entity recognition**: [bert_ner](/natural_language_processing/bert_ner/), [t5_base_japanese_ner](/natural_language_processing/t5_base_japanese_ner/), [bert_ner_japanese](/natural_language_processing/bert_ner_japanese/)
**Reranker**: [cross_encoder_mmarco](/natural_language_processing/cross_encoder_mmarco/), [japanese-reranker-cross-encoder](/natural_language_processing/japanese-reranker-cross-encoder/), [ruri-v3-reranker](/natural_language_processing/ruri-v3-reranker/)
**Sentence generation**: [gpt2](/natural_language_processing/gpt2/), [rinna_gpt2](/natural_language_processing/rinna_gpt2/)
**Sentiment analysis**: [bert_sentiment_analysis](/natural_language_processing/bert_sentiment_analysis/), [bert_tweets_sentiment](/natural_language_processing/bert_tweets_sentiment/)
**Summarize**: [bert_sum_ext](/natural_language_processing/bert_sum_ext/), [presumm](/natural_language_processing/presumm/), [t5_base_japanese_title_generation](/natural_language_processing/t5_base_japanese_title_generation/), [t5_base_summarization](/natural_language_processing/t5_base_japanese_summarization/)
**Translation**: [fugumt-en-ja](/natural_language_processing/fugumt-en-ja/), [fugumt-ja-en](/natural_language_processing/fugumt-ja-en/)
**Zero shot classification**: [bert_zero_shot_classification](/natural_language_processing/bert_zero_shot_classification/), [multilingual-minilmv2](/natural_language_processing/multilingual-minilmv2/) |
| | [Network intrusion detection](/network_intrusion_detection/) | [bert-network-packet-flow-header-payload](/network_intrusion_detection/bert-network-packet-flow-header-payload/), [falcon-adapter-network-packet](/network_intrusion_detection/falcon-adapter-network-packet/) |
| [
](/neural_rendering/) | [Neural rendering](/neural_rendering/) | [nerf](/neural_rendering/nerf/), [TripoSR](/neural_rendering/tripo_sr/) |
| | [NSFW detector](/nsfw_detector/) | [clip-based-nsfw-detector](/nsfw_detector/clip-based-nsfw-detector/) |
| [
](/object_detection/) | [Object detection](/object_detection/) | **CNN**: [yolov1-tiny](/object_detection/yolov1-tiny/), [yolov2](/object_detection/yolov2/), [yolov2-tiny](/object_detection/yolov2-tiny/), [maskrcnn](/object_detection/maskrcnn/), [yolov3](/object_detection/yolov3/), [yolov3-tiny](/object_detection/yolov3-tiny/), [mobilenet_ssd](/object_detection/mobilenet_ssd/), [m2det](/object_detection/m2det/), [centernet](/object_detection/centernet/), [yolact](/object_detection/yolact/), [efficientdet](/object_detection/efficientdet/), [pedestrian_detection](/object_detection/pedestrian_detection/), [crowd_det](/object_detection/crowd_det/), [yolov4](/object_detection/yolov4/), [yolov4-tiny](/object_detection/yolov4-tiny/), [yolov5](/object_detection/yolov5/), [poly_yolo](/object_detection/poly_yolo/), [nanodet](/object_detection/nanodet/), [yolor](/object_detection/yolor/), [yolox](/object_detection/yolox/), [picodet](/object_detection/picodet/), [yolox-ti-lite](/object_detection/yolox-ti-lite/), [yolov7](/object_detection/yolov7/), [fastest-det](/object_detection/fastest-det/), [yolov](/object_detection/yolov/), [yolov6](/object_detection/yolov6/), [damo_yolo](/object_detection/damo_yolo/), [yolov8](/object_detection/yolov8/), [yolox_body_head_hand_face](/object_detection/yolox_body_head_hand_face/), [yolov9](/object_detection/yolov9/), [yolov10](/object_detection/yolov10/), [yolov11](/object_detection/yolov11/), [yolov12](/object_detection/yolov12/)
**Transformer**: [detr](/object_detection/detr/), [glip](/object_detection/glip/), [dab-detr](/object_detection/dab-detr/), [detic](/object_detection/detic/), [groundingdino](/object_detection/groundingdino/), [rt-detr-v2](/object_detection/rt-detr-v2/)
**Specific target**: [traffic-sign-detection](/object_detection/traffic-sign-detection/), [sku110k-densedet](/object_detection/sku110k-densedet/), [footandball](/object_detection/footandball/), [qrcode_wechatqrcode](/object_detection/qrcode_wechatqrcode/), [mobile_object_localizer](/object_detection/mobile_object_localizer/), [layout_parsing](/object_detection/layout_parsing/) |
| [
](/object_detection_3d/) | [Object detection 3d](/object_detection_3d/) | [3d_bbox](/object_detection_3d/3d_bbox/), [d4lcn](/object_detection_3d/d4lcn/), [egonet](/object_detection_3d/egonet/), [mediapipe_objectron](/object_detection_3d/mediapipe_objectron/), [3d-object-detection.pytorch](/object_detection_3d/3d-object-detection.pytorch/), [did_m3d](/object_detection_3d/did_m3d/) |
| [
](/object_tracking/) | [Object tracking](/object_tracking/) | [deepsort](/object_tracking/deepsort/), [person_reid_baseline_pytorch](/object_tracking/person_reid_baseline_pytorch/), [abd_net](/object_tracking/abd_net/), [deepsort_vehicle](/object_tracking/deepsort_vehicle/), [qd-3dt](/object_tracking/qd-3dt/), [centroids-reid](/object_tracking/centroids-reid/), [siam-mot](/object_tracking/siam-mot/), [bytetrack](/object_tracking/bytetrack/), [strong_sort](/object_tracking/strong_sort/), [samurai](/object_tracking/samurai/) |
| [
](/optical_flow_estimation/) | [Optical flow estimation](/optical_flow_estimation/) | [raft](/optical_flow_estimation/raft/), [cotracker3](/optical_flow_estimation/cotracker3/) |
| [
](/point_segmentation/) | [Point segmentation](/point_segmentation/) | [pointnet_pytorch](/point_segmentation/pointnet_pytorch/) |
| [
](/pose_estimation/) | [Pose estimation](/pose_estimation/) | [openpose](/pose_estimation/openpose/), [posenet](/pose_estimation/posenet/), [pose_resnet](/pose_estimation/pose_resnet/), [lightweight-human-pose-estimation](/pose_estimation/lightweight-human-pose-estimation/), [animalpose](/pose_estimation/animalpose/), [efficientpose](/pose_estimation/efficientpose/), [blazepose](/pose_estimation/blazepose/), [mediapipe_holistic](/pose_estimation/mediapipe_holistic/), [movenet](/pose_estimation/movenet/), [ap-10k](/pose_estimation/ap-10k/), [e2pose](/pose_estimation/e2pose/) |
| [
](/pose_estimation_3d/) | [Pose estimation 3d](/pose_estimation_3d/) | [pose-hg-3d](/pose_estimation_3d/pose-hg-3d/), [3d-pose-baseline](/pose_estimation_3d/3d-pose-baseline/), [lightweight-human-pose-estimation-3d](/pose_estimation_3d/lightweight-human-pose-estimation-3d/), [3dmppe_posenet](/pose_estimation_3d/3dmppe_posenet/), [gast](/pose_estimation_3d/gast/), [blazepose-fullbody](/pose_estimation_3d/blazepose-fullbody/), [mediapipe_pose_world_landmarks](/pose_estimation_3d/mediapipe_pose_world_landmarks/) |
| [
](/road_detection/) | [Road detection](/road_detection/) | [road-segmentation-adas](/road_detection/road-segmentation-adas/), [codes-for-lane-detection](/road_detection/codes-for-lane-detection/), [ultra-fast-lane-detection](/road_detection/ultra-fast-lane-detection/), [polylanenet](/road_detection/polylanenet/), [roneld](/road_detection/roneld/), [lstr](/road_detection/lstr/), [yolop](/road_detection/yolop/), [cdnet](/road_detection/cdnet/), [hybridnets](/road_detection/hybridnets/) |
| [
](/rotation_prediction/) | [Rotation prediction](/rotation_prediction/) | [rotnet](/rotation_prediction/rotnet/) |
| [
](/style_transfer/) | [Style transfer](/style_transfer/) | [adain](/style_transfer/adain/), [pix2pixHD](/style_transfer/pix2pixHD/), [beauty_gan](/style_transfer/beauty_gan/), [psgan](/style_transfer/psgan/), [animeganv2](/style_transfer/animeganv2/), [EleGANt](/style_transfer/elegant/) |
| [
](/super_resolution/) | [Super resolution](/super_resolution/) | [srresnet](/super_resolution/srresnet/), [edsr](/super_resolution/edsr/), [han](/super_resolution/han/), [real-esrgan](/super_resolution/real-esrgan/), [swinir](/super_resolution/swinir/), [rcan-it](/super_resolution/rcan-it/), [Hat](/super_resolution/hat/), [SPAN](/super_resolution/span/) |
| [
](/text_detection/) | [Text detection](/text_detection/) | [east](/text_detection/east/), [pixel_link](/text_detection/pixel_link/), [craft_pytorch](/text_detection/craft_pytorch/) |
| [
](/text_recognition/) | [Text recognition](/text_recognition/) | [etl](/text_recognition/etl/), [crnn.pytorch](/text_recognition/crnn.pytorch/), [deep-text-recognition-benchmark](/text_recognition/deep-text-recognition-benchmark/), [easyocr](/text_recognition/easyocr/), [paddleocr](/text_recognition/paddleocr/), [donut](/text_recognition/donut/), [ndlocr_text_recognition](/text_recognition/ndlocr_text_recognition/), [paddleocr_v3](/text_recognition/paddleocr_v3/) |
| | [Time-series forecasting](/time_series_forecasting/) | [informer2020](/time_series_forecasting/informer2020/), [timesfm](/time_series_forecasting/timesfm/), [moirai](/time_series_forecasting/moirai/), [chronos2](/time_series_forecasting/chronos2/) |
| [
](/vehicle_recognition/) | [Vehicle recognition](/vehicle_recognition/) | [vehicle-attributes-recognition-barrier](/vehicle_recognition/vehicle-attributes-recognition-barrier/), [vehicle-license-plate-detection-barrier](/vehicle_recognition/vehicle-license-plate-detection-barrier/) |
| [
](/vision_language_model/) | [Vision language model](/vision_language_model/) | [llava](/vision_language_model/llava/), [florence2](/vision_language_model/florence2/), [mobilevlm](/vision_language_model/mobilevlm/), [llava-jp](/vision_language_model/llava-jp/), [qwen2_vl](/vision_language_model/qwen2_vl/), [qwen2.5_vl](/vision_language_model/qwen2.5_vl/), [qwen3_vl](/vision_language_model/qwen3_vl/) |
| | [Commercial model](/commercial_model/) | [acculus-pose](/commercial_model/acculus-pose/) |
# About ailia SDK
[ailia SDK](https://ailia.ai/en/sdk/) is a cross-platform, high-speed inference SDK for AI. It supports Windows, Mac, Linux, iOS, Android, Jetson, and Raspberry Pi with GPU acceleration via Vulkan and Metal. Bindings are available for C++, Python, Unity (C#), Kotlin, Rust, and Flutter.
# Other platforms
Prototype with ailia MODELS (Python), then deploy to production.
- [unity version](https://github.com/ailia-ai/ailia-models-unity)
- [kotlin version](https://github.com/ailia-ai/ailia-models-kotlin)
- [c++ version](https://github.com/ailia-ai/ailia-models-cpp)
- [flutter version](https://github.com/ailia-ai/ailia-models-flutter)
- [rust version](https://github.com/ailia-ai/ailia-models-rust)
- [js version](https://github.com/ailia-ai/ailia-models-js)
# Contact
- [Contact us](https://www.ailia.ai/en-contact-product)
- [Mail](mailto:contact@ailia.ai)