KIE.AI
All Model
  • All Model
  • Old Model
language
language
  • 🇺🇸 English
  • 🇨🇳 Chinese
language
language
  • 🇺🇸 English
  • 🇨🇳 Chinese
Support
All Model
  • All Model
  • Old Model
All Model
  • All Model
  • Old Model
Market
File Upload APICommon API
Market
File Upload APICommon API
  1. Vocal Removal
  • Getting Started with KIE API (Important)
  • Market
  • Image Models
    • Seedream
      • Seedream4.0 - Text to Image
      • Seedream4.0 - Edit
      • Seedream4.5 - Text to Image
      • Seedream4.5 - Edit
      • Seedream5.0 Lite - Text to Image
      • Seedream5.0 Lite - Image to Image
      • Seedream5.0 Pro - Text to Image
      • Seedream5.0 Pro - Image to Image
      • Seedream 5.0 Pro - Layer Decomposition
      • Seedream3.0 - Text to Image
    • Z-image
      • Z-Image
    • Google
      • Google - imagen4-fast
      • Google - imagen4-ultra
      • Google - imagen4
      • Google - Nano Banana Edit
      • Google - Nano Banana
      • Google - Nano Banana Pro
      • Google - Nano Banana 2
      • Google - Nano Banana 2 Lite
    • Flux-2
      • Flux-2 - Pro Image to Image
      • Flux-2 - Pro Text to Image
      • Flux-2 - Image to Image
      • Flux-2 - Text to Image
    • Grok Imagine
      • Grok Imagine Image 2.0 Text To Image
      • Grok Imagine Image 2.0 Segment Map
      • Grok Imagine Image 2.0 Segment Edit
      • Grok Imagine - Text to Image
      • Grok Imagine Image 2.0 Image Edit
      • Grok Imagine - image to image
    • GPT Image
      • GPT Image 2.5 Flare - Text to Image
      • GPT Image 2.5 Flare - Image To Image
      • GPT Image 2.5 Sunburst - Text to Image
      • GPT Image 2.5 Sunburst - Image To Image
      • GPT Image-1.5 - Text to Image
      • GPT Image-1.5 - Image to Image
      • GPT Image-2 - Text to Image
      • GPT Image 2 - Image To Image
    • Topaz
      • Topaz - Image Upscale
    • Recraft
      • Recraft - Remove Background
      • Recraft - Crisp Upscale
    • Ideogram
      • Ideogram - Character Edit
      • Ideogram - Character Remix
      • Ideogram - Character
      • Ideogram V3 Text to Image
      • Ideogram V3 Edit
      • Ideogram V3 Remix
    • Qwen
      • Qwen - Text to Image
      • Qwen - Image to Image
      • Qwen - Image Edit
      • Qwen2 - Image Edit
      • Qwen2 - Text To Image
      • Qwen3 Pro Text to Image
      • Qwen3 Text to Image
      • Qwen3 Pro Image to Image
      • Qwen3 Image to Image
    • Wan
      • Wan 2.7 Image
      • Wan 2.7 Image Pro
    • 4o Image API
      • 4o Image Generation Callbacks
      • Generate 4o Image
    • Flux Kontext API
      • Image Generation or Editing Callbacks
      • Generate or Edit Image
  • Video Models
    • Grok Imagine
      • Grok Imagine Text to Video
      • Grok Imagine Image to Video
      • Grok Imagine - Video Upscale
      • Grok Imagine - Video Extend
      • Grok Imagine Video 1.5 Preview
    • Kling
      • Kling 2.6 Text to Video
      • Kling 2.6 Image to Video
      • Kling - V2.5 Turbo Image to Video Pro
      • Kling - V2.5 Turbo Text to Video Pro
      • Kling AI Avatar Standard
      • Kling AI Avatar Pro
      • Kling V2.1 Master Image to Video
      • Kling V2.1 Master Text to Video
      • Kling V2.1 Pro
      • Kling V2.1 Standard
      • Kling 2.6 motion-control
      • Kling-3.0 motion-control
      • Kling 3.0
      • Kling - V3 Turbo Text to Video
      • Kling - V3 Turbo Image to Video
      • Kling 3.0 Omni Reference To Video
      • Kling 3.0 Omni Transformation
      • Kling 3.0 Omni Image To Video
      • Kling 3.0 Omni Text to Video
    • Bytedance
      • Bytedance Seedance 2.0
      • Bytedance Seedance 2.0 Fast
      • Bytedance Seedance 2.0 Mini
      • Bytedance Seedance 2.5
      • Bytedance Seedance 1.5 Pro
      • Bytedance V1 Pro Fast Image to Video
      • Bytedance V1 Pro Image to Video
      • Bytedance - V1 Pro Text to Video
      • Bytedance - V1 Lite Image to Video
      • Bytedance - V1 Lite Text to Video
    • Hailuo
      • Hailuo 2.3 Pro Image to Video
      • Hailuo 2.3 Standard Image to Video
      • Hailuo Pro Text to Video
      • Hailuo Pro Image to Video
      • Hailuo Standard Text to Video
      • Hailuo Standard Image to Video
    • Wan
      • Wan - 2.2 A14B Image to Video Turbo
      • Wan - 2.2 A14B Speech to Video Turbo
      • Wan - 2.2 A14B Text to Video Turbo
      • Wan - Animate Move
      • Wan - Animate Replace
      • Wan 2.6 - Image to Video
      • Wan 2.6 - Text to Video
      • Wan 2.6 - Video to Video
      • Wan - 2.6-flash-image-to-video
      • Wan - 2-6-flash-video-to-video
      • Wan 2.5 - Image to Video
      • Wan 2.5 - Text to Video
      • Wan 2.7 - Text to Video
      • Wan 2.7 - Image to Video
      • Wan 2.7 - Video Edit
      • Wan 2.7 - Reference to Video
      • Wan 3.0 - Video
      • Wan 3.0 - Video Prime
    • Topaz
      • Topaz - Video Upscale
    • Infinitalk
      • Infinitalk - From Audio
    • PixVerse
      • PixVerse V6 Text-to-Video
      • PixVerse V6 Image-to-Video
      • PixVerse V6 First & Last Frame Transition
      • PixVerse V6 Video Extension
      • PixVerse V6 Fusion / Reference-to-Video
    • MiniMax H3
      • MiniMax H3 Text-to-Video
      • MiniMax H3 Image-to-Video
      • MiniMax H3 Reference-to-Video
    • Runway API
      • AI Video Generation Callbacks
      • AI Video Extension Callbacks
      • Aleph
        • Aleph Video Generation Callbacks
        • Generate Aleph Video
      • Generate AI Video
      • Extend AI Video
    • HappyHorse
      • HappyHorse - text-to-video
      • HappyHorse - image-to-video
      • HappyHorse - reference-to-video
      • HappyHorse - video-edit
      • HappyHorse-1-1 image-to-video
      • HappyHorse-1-1 text-to-video
      • HappyHorse-1-1 reference-to-video
    • Gemini Omni
      • Gemini Omni 1.1 Flash
      • Gemini Omni Video
      • Gemini Omni Audio
      • Gemini Omni Character
    • OmniHuman
      • Omnihuman 1.5
      • Omnihuman 1.5 Human Identification
      • OmniHuman 1.5 Subject Detection
    • Volcengine
      • Volcengine video to video lip sync
    • Veo3.1 API
      • Veo3.1 Video Generation Callbacks
      • Get 4K Video Callbacks
      • Generate Veo3.1 Video
      • Get 1080P Video
      • Get 4K Video
      • Extend Veo3.1 Video
  • Music Models
    • ElevenLabs
      • elevenlabs/audio-isolation
      • elevenlabs/text-to-dialogue-v3
      • elevenlabs/text-to-speech-multilingual-v2
      • elevenlabs/text-to-speech-turbo-2-5
    • Gemini
      • Gemini 3.1 Flash Text to speech
      • Gemini 2.5 Pro Text to Speech
    • Suno
      • Music Generation
        • Music Cover Generation Callbacks
        • Music Generation Callbacks
        • Music Extension Callbacks
        • Audio Upload and Cover Callbacks
        • Audio Upload and Extension Callbacks
        • Add Instrumental Callbacks
        • Add Vocals Callbacks
        • Replace Music Section Callbacks
        • Generate Music
        • Extend Music
        • Upload And Cover Audio
        • Upload And Extend Audio
        • Add Instrumental to Music
        • Add Vocals to Music
        • Get Timestamped Lyrics
        • Boost Music Style
        • Generate Music Cover
        • Replace Music Section
        • Generate Persona
        • Generate Mashup Music
        • Recovery Audio
      • WAV Conversion
        • Convert to WAV Callbacks
        • Convert to WAV Format
      • Music Video Generation
        • Music Video Generation Callbacks
        • Create Music Video
      • Lyrics Generation
        • Lyrics Generation Callbacks
        • Generate Lyrics
      • voice
        • Suno Voice Generation Callback
        • Suno Voice Validation Phrase Callback
        • Suno Voice Generate Verification Phrase API
        • Suno Voice Create Custom Voice API
        • Suno Voice Regenerate Verification Phrase
        • Suno Voice Check Availability API
      • Vocal Removal
        • Audio Separation Callbacks
        • MIDI Generation Callbacks
        • Vocal & Instrument Stem Separation
          POST
        • Generate MIDI from Audio
          POST
      • Sounds Generation
        • Generate sounds
  • Chat Models
    • GPT
      • GPT 5.2
      • GPT 5.4 (response)
      • GPT 5.5 (response)
      • GPT 5.6 Luna
      • Gpt 6 Astra
      • GPT 5.6 Terra
      • GPT 5.6 Sol
    • Claude
      • Claude Code + kie.ai Integration Guide
      • Claude Opus 4.7
      • Claude Opus 4.8
      • Claude Fable 5
      • Claude Sonnet 5
      • Claude Haiku 4.5
      • Claude Opus 4.5
      • Claude Opus 4.6
      • Claude Opus 5
      • Claude Sonnet 4.5
      • Claude Sonnet 4.6
    • Codex
      • GPT Codex
    • Gemini
      • Gemini 2.5 Pro (openai)
      • Gemini 3 Pro (openai)
      • Gemini 3.1 Pro (openai)
      • Gemini 2.5 Flash (openai)
      • Gemini 3 Flash (openai)
      • Gemini 3.5 Flash
      • Gemini 3.5 Flash (openai)
      • Gemini 3.6 Flash
      • Gemini 3.6 Flash (openai)
      • Gemini 3.7 Flash
      • Gemini 3.7 Flash (openai)
      • Gemini 3.8 Flash
      • Gemini 3.8 Flash (openai)
      • Gemini 3 Flash
    • Grok
      • Grok 4.3
      • Grok 4.5
      • Grok 4.6
  • Get Task Details
    GET
  1. Vocal Removal

Generate MIDI from Audio

POST
/api/v1/jobs/createTask
Document updated
If you have already completed the integration process previously, you can still access the old version document at: Old version address (https://docs.kie.ai/old-model/suno-api/generate-midi)
Convert separated audio tracks into MIDI format with detailed note information for each instrument.

Usage Guide#

Convert separated audio tracks into structured MIDI data containing pitch, timing, and velocity information
Requires a completed vocal separation task ID (from the Vocal Removal API)
Generates MIDI note data for multiple detected instruments including drums, bass, guitar, keyboards, and more
Ideal for music transcription, notation, remixing, or educational analysis
Best results on clean, well-separated audio tracks with clear instrument parts

Prerequisites#

Required
You must first use the Vocal & Instrument Stem Separation API to separate your audio before generating MIDI.

Parameter Reference#

NameTypeDescription
task_idstringRequired. Task ID from a completed vocal separation.
callBackUrlstringRequired. URL to receive MIDI generation completion notifications.
audio_idstringOptional. Specifies which separated audio track to generate MIDI from. This audio_id can be obtained from the originData array in the Get Vocal Separation Details endpoint response. Each item in originData contains an id field that can be used here. If not provided, MIDI will be generated from all separated tracks.

Developer Notes#

The callBackUrl will contain detailed note data for each detected instrument.
Each note includes: pitch (MIDI note number), start (seconds), end (seconds), velocity (0-1).
Not all instruments may be detected — depends on audio content.
Pricing: Check current per-call credit costs at https://kie.ai/pricing.

Request

Authorization
Bearer Token
Provide your bearer token in the
Authorization
header when making requests to protected resources.
Example:
Authorization: Bearer ********************
or
Body Params application/jsonRequired

Examples

Responses

🟢200
application/json
MIDI generation task created successfully
Bodyapplication/json

🔴500Error
Request Request Example
Shell
JavaScript
Java
Swift
curl --location 'https://api.kie.ai/api/v1/jobs/createTask' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '{
  "model": "ai-music-api/generate-midi-from-audio",
  "callBackUrl": "https://example.callback",
  "input": {
    "task_id": "5c79****be8e",
    "audio_id": "8ca376e7-******-08aaf2c6dd27"
  }
}'
Response Response Example
200 - 成功示例
{
    "code": 200,
    "msg": "success",
    "data": {
        "taskId": "5c79****be8e"
    }
}
Previous
Vocal & Instrument Stem Separation
Next
Generate sounds
Built with