Year Entrytype Title Author Link Code Task Reproducible Dataset Framework Architecture Dropout Batch Epochs Dataaugmentation Input Dimension Activation Loss Learningrate Optimizer Gpu 1988 inproceedings Neural net modeling of music Bharucha, J. 1988 inproceedings Creation by refinement: A creativity paradigm for gradient descent learning networks Lewis, J. P. http://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=23933 1988 inproceedings A sequential network design for musical applications Todd, Peter M. Composition 1989 article The representation of pitch in a neural net model of chord classification Laden, Bernice and Keefe, Douglas H. http://www.jstor.org/stable/3679550 Chord recognition 1989 inproceedings Algorithms for music composition by neural nets: Improved CBR paradigms Lewis, J. P. https://quod.lib.umich.edu/cgi/p/pod/dod-idx/algorithms-for-music-composition.pdf?c=icmc;idno=bbp2372.1989.044;format=pdf Composition 1989 article A connectionist approach to algorithmic composition Todd, Peter M. http://www.jstor.org/stable/3679551 Composition 1994 article Neural network music composition by prediction: Exploring the benefits of psychoacoustic constraints and multi-scale processing Mozer, Michael C. http://www-labs.iro.umontreal.ca/~pift6080/H09/documents/papers/mozer-music.pdf Composition 1995 inproceedings Automatic source identification of monophonic musical instrument sounds Kaminsky, I. and Materka, Andrzej https://www.researchgate.net/publication/3622871_Automatic_source_identification_of_monophonic_musical_instrument_sounds No Instrument recognition No Inhouse No No No No No No Raw audio 1D Sigmoid No 0.25 No No 1995 inproceedings Neural network based model for classification of music type Matityaho, Benyamin and Furst, Miriam http://ieeexplore.ieee.org/abstract/document/514161/ MGR 1997 inproceedings A machine learning approach to musical style recognition Dannenberg, Roger B and Thom, Belinda and Watson, David http://repository.cmu.edu/cgi/viewcontent.cgi?article=1496&context=compsci MSR 1998 inproceedings Recognition of music types Soltau, Hagen and Schultz, Tanja and Westphal, Martin and Waibel, Alex https://www.ri.cmu.edu/pub_files/pub1/soltau_hagen_1998_2/soltau_hagen_1998_2.pdf No MGR No Inhouse No DNN No No No No 10x5 cepstral coefficients 2D No No No No No 1999 book Musical networks: Parallel distributed perception and performance Griffith, Niall and Todd, Peter M. https://s3.amazonaws.com/academia.edu.documents/3551783/10.1.1.39.6248.pdf?AWSAccessKeyId=AKIAIWOWYYGZ2Y53UL3A&Expires=1507055806&Signature=5mGzQc7bvJgUZYfXOmCX8eeNQOs%3D&response-content-disposition=inline%3B%20filename%3DMusical_networks_Parallel_distributed_pe.pdf 2001 inproceedings Multi-phase learning for jazz improvisation and interaction Franklin, Judy A http://www.cs.smith.edu/~jfrankli/papers/CtColl01.pdf Composition RNN 2002 inproceedings A supervised learning approach to musical style recognition Buzzanca, Giuseppe https://www.researchgate.net/profile/Giuseppe_Buzzanca/publication/228588086_A_supervised_learning_approach_to_musical_style_recognition/links/54b43ee90cf26833efd0109f.pdf MGR 2002 inproceedings Finding temporal structure in music: Blues improvisation with LSTM recurrent networks Eck, Douglas and Schmidhuber, Juergen http://www-perso.iro.umontreal.ca/~eckdoug/papers/2002_ieee.pdf No Composition No Inhouse No RNN-LSTM No No No No Midi Chords & Midi notes 1D Logistic Sigmoid cross-entropy 0.00001 SGD No 2002 unpublished Neural networks for note onset detection in piano music Marolt, Matija and Kavcic, Alenka and Privosnik, Marko https://www.researchgate.net/profile/Matija_Marolt/publication/2473938_Neural_Networks_for_Note_Onset_Detection_in_Piano_Music/links/00b49525efccc79fed000000.pdf No Onset detection No Inhouse No MLP No No No No Raw audio signal and synthesized 1D No No No No No 2004 inproceedings A convolutional-kernel based approach for note onset detection in piano-solo audio signals Nava, Gabriel Pablo and Tanaka, Hidehiko and Ide, Ichiro http://www.murase.nuie.nagoya-u.ac.jp/~ide/res/paper/E04-conference-pablo-1.pdf Onset detection 2009 inproceedings Unsupervised feature learning for audio classification using convolutional deep belief networks Lee, Honglak and Pham, Peter and Largman, Yan and Ng, Andrew Y http://papers.nips.cc/paper/3674-unsupervised-feature-learning-for-audio-classification-using-convolutional-deep-belief-networks.pdf Speaker gender recognition [TIMIT](https://catalog.ldc.upenn.edu/LDC93S1) CDBN 2010 phdthesis Audio musical genre classification using convolutional neural networks and pitch and tempo transformations Li, Lihua http://lbms03.cityu.edu.hk/theses/c_ftt/mphil-cs-b39478026f.pdf [GTzan](http://marsyas.info/downloads/datasets.html) MFCC 2010 inproceedings Automatic musical pattern feature extraction using convolutional neural network Li, Tom LH and Chan, Antoni B and Chun, A https://www.researchgate.net/profile/Antoni_Chan2/publication/44260643_Automatic_Musical_Pattern_Feature_Extraction_Using_Convolutional_Neural_Network/links/02e7e523dac6bb86b0000000.pdf MGR [GTzan](http://marsyas.info/downloads/datasets.html) MFCC 2011 inproceedings Audio-based music classification with a pretrained convolutional network Dieleman, Sander and Brakel, Philémon and Schrauwen, Benjamin http://www.ismir2011.ismir.net/papers/PS6-3.pdf No MGR & Artist recognition No [MSD](https://labrosa.ee.columbia.edu/millionsong/) Theano CNN & MLP 0.3 No 1 No Custom 0.005 & 0.0001 No No 2012 inproceedings Rethinking automatic chord recognition with convolutional neural networks Humphrey, Eric J. and Bello, Juan Pablo http://ieeexplore.ieee.org/abstract/document/6406762/ Chord recognition [Beatles](http://isophonics.net/content/reference-annotations-beatles) & [RWC](https://staff.aist.go.jp/m.goto/RWC-MDB/) & [US Pop](https://labrosa.ee.columbia.edu/projects/musicsim/uspop2002.html) CNN Cross-entropy 2012 inproceedings Moving beyond feature design: Deep architectures and automatic feature learning in music informatics Humphrey, Eric J. and Bello, Juan Pablo and LeCun, Yann http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.294.2304&rep=rep1&type=pdf 2012 inproceedings Local-feature-map integration using convolutional neural networks for music genre classification Nakashika, Toru and Garcia, Christophe and Takiguchi, Tetsuya and De Lyon, Insa http://liris.cnrs.fr/Documents/Liris-5602.pdf MGR [GTzan](http://marsyas.info/downloads/datasets.html) GLCM 2012 inproceedings Learning sparse feature representations for music annotation and retrieval Nam, Juhan and Herrera, Jorge and Slaney, Malcolm and Smith, Julius O https://pdfs.semanticscholar.org/099d/85f25e9336f48ff64287a4b53ee5fb64ab51.pdf 2012 inproceedings Unsupervised learning of local features for music classification Wülfing, Jan and Riedmiller, Martin http://www.ismir2012.ismir.net/event/papers/139_ISMIR_2012.pdf MGR [GTzan](http://marsyas.info/downloads/datasets.html) CQT 2013 inproceedings Multiscale approaches to music audio feature learning Dieleman, Sander and Schrauwen, Benjamin http://ismir2013.ismir.net/wp-content/uploads/2013/09/69_Paper.pdf [Magnatagatune](http://mirg.city.ac.uk/codeapps/the-magnatagatune-dataset) Mel-spectrogram cross-entropy 2013 inproceedings Musical onset detection with convolutional neural networks Schlüter, Jan and Böck, Sebastian http://phenicx.upf.edu/system/files/publications/Schlueter_MML13.pdf Onset detection Mel-spectrogram cross-entropy 2013 inproceedings Deep content-based music recommendation Van den Oord, Aaron and Dieleman, Sander and Schrauwen, Benjamin http://papers.nips.cc/paper/5004-deep-content-based-music-recommendation.pdf Recommendation [MSD](https://labrosa.ee.columbia.edu/millionsong/) & [Echo Nest Taste Profile Subset](https://labrosa.ee.columbia.edu/millionsong/tasteprofile) & [Last.fm](https://www.last.fm/) Theano CNN MFCC & Mel-Spectro ReLU 2014 inproceedings The munich LSTM-RNN approach to the MediaEval 2014 Emotion In Music task Coutinho, Eduardo and Weninger, Felix and Schuller, Björn W and Scherer, Klaus R https://pdfs.semanticscholar.org/8a24/c5131d5a28165f719697028c34b00e6d3f60.pdf MER RNN-LSTM 2014 inproceedings End-to-end learning for music audio Dieleman, Sander and Schrauwen, Benjamin http://ieeexplore.ieee.org/abstract/document/6854950/ MGR [Magnatagatune](http://mirg.city.ac.uk/codeapps/the-magnatagatune-dataset) CNN Raw & Mel-spectrogram 2014 techreport Deep learning for music genre classification Feng, Tao https://courses.engr.illinois.edu/ece544na/fa2014/Tao_Feng.pdf MGR [GTzan](http://marsyas.info/downloads/datasets.html) 2014 inproceedings Recognition of acoustic events using deep neural networks Gencoglu, Oguzhan and Virtanen, Tuomas and Huttunen, Heikki https://www.cs.tut.fi/sgn/arg/music/tuomasv/dnn_eusipco2014.pdf 2014 article Deep image features in music information retrieval Gwardys, Grzegorz and Grzywczak, Daniel https://www.degruyter.com/downloadpdf/j/eletel.2014.60.issue-4/eletel-2014-0042/eletel-2014-0042.pdf [GTzan](http://marsyas.info/downloads/datasets.html) 2014 inproceedings From music audio to chord tablature: Teaching deep convolutional networks to play guitar Humphrey, Eric J. and Bello, Juan Pablo https://ejhumphrey.com/assets/pdf/humphrey2014music.pdf Chord recognition [Beatles](http://isophonics.net/content/reference-annotations-beatles) & [RWC](https://staff.aist.go.jp/m.goto/RWC-MDB/) & [US Pop](https://labrosa.ee.columbia.edu/projects/musicsim/uspop2002.html) CQT 2014 inproceedings Improved musical onset detection with convolutional neural networks Schlüter, Jan and Bock, Sebastian http://www.mirlab.org/conference_papers/International_Conference/ICASSP%202014/papers/p7029-schluter.pdf Onset detection Inhouse CNN Mel-spectrogram 3D 2014 inproceedings Boundary detection in music structure analysis using convolutional neural networks Ullrich, Karen and Schlüter, Jan and Grill, Thomas https://dav.grrrr.org/public/pub/ullrich_schlueter_grill-2014-ismir.pdf Boundary detection [SALAMI](http://ddmal.music.mcgill.ca/research/salami/annotations) Mel-spectrogram Cross-entropy 2014 inproceedings Improving content-based and hybrid music recommendation using deep learning Wang, Xinxi and Wang, Ye http://www.smcnus.org/wp-content/uploads/2014/08/reco_MM14.pdf Recommendation [Echo Nest Taste Profile Subset](https://labrosa.ee.columbia.edu/millionsong/tasteprofile) & [7digital](https://7digital.com) Theano DBN No 15 nodes of 2 Tesla M2090 2014 inproceedings A deep representation for invariance and music classification Zhang, Chiyuan and Evangelopoulos, Georgios and Voinea, Stephen and Rosasco, Lorenzo and Poggio, Tomaso http://www.mirlab.org/conference_papers/International_Conference/ICASSP%202014/papers/p7034-zhang.pdf MGR [GTzan](http://marsyas.info/downloads/datasets.html) CNN 2015 inproceedings Auralisation of deep convolutional neural networks: Listening to learned features Choi, Keunwoo and Fazekas, György and Sandler, Mark Brian and Kim, Jeonghee http://ismir2015.uma.es/LBD/LBD24.pdf https://github.com/keunwoochoi/Auralisation MGR Inhouse STFT 2015 inproceedings Downbeat tracking with multiple features and deep neural networks Durand, Simon and Bello, Juan Pablo and David, Bertrand and Richard, Gaël http://perso.telecom-paristech.fr/~grichard/Publications/2015-durand-icassp.pdf Beat detection 2015 inproceedings Music boundary detection using neural networks on spectrograms and self-similarity lag matrices Grill, Thomas and Schlüter, Jan http://www.ofai.at/~jan.schlueter/pubs/2015_eusipco.pdf Boundary detection [SALAMI](http://ddmal.music.mcgill.ca/research/salami/annotations) STFT 2015 inproceedings Classification of spatial audio location and content using convolutional neural networks Hirvonen, Toni https://www.researchgate.net/profile/Toni_Hirvonen/publication/276061831_Classification_of_Spatial_Audio_Location_and_Content_Using_Convolutional_Neural_Networks/links/5550665908ae12808b37fe5a/Classification-of-Spatial-Audio-Location-and-Content-Using-Convolutional-Neural-Networks.pdf 2015 inproceedings Deep learning, audio adversaries, and music content analysis Kereliuk, Corey and Sturm, Bob L. and Larsen, Jan http://www2.imm.dtu.dk/pubdb/views/edoc_download.php/6905/pdf/imm6905.pdf 2015 article Deep learning and music adversaries Kereliuk, Corey and Sturm, Bob L. and Larsen, Jan https://arxiv.org/pdf/1507.04761.pdf https://github.com/coreyker/dnn-mgr MGR [GTzan](http://marsyas.info/downloads/datasets.html) & [LMD](https://sites.google.com/site/carlossillajr/resources/the-latin-music-database-lmd) CNN Magnitude spectral frames 2015 inproceedings Singing voice detection with deep recurrent neural networks Leglaive, Simon and Hennequin, Romain and Badeau, Roland https://hal-imt.archives-ouvertes.fr/hal-01110035/ SVD 2015 unpublished Automatic instrument recognition in polyphonic music using convolutional neural networks Li, Peter and Qian, Jiyuan and Wang, Tian https://arxiv.org/pdf/1511.05520.pdf Instrument recognition [MedleyDB](http://medleydb.weebly.com/) 1D Freq Raw audio Cross-entropy 2015 inproceedings A software framework for musical data augmentation McFee, Brian and Humphrey, Eric J. and Bello, Juan Pablo https://bmcfee.github.io/papers/ismir2015_augmentation.pdf Instrument recognition [MedleyDB](http://medleydb.weebly.com/) 2015 unpublished A deep bag-of-features model for music auto-tagging Nam, Juhan and Herrera, Jorge and Lee, Kyogu https://arxiv.org/pdf/1508.04999v1.pdf 2015 inproceedings Music-noise segmentation in spectrotemporal domain using convolutional neural networks Park, Taejin and Lee, Taejin http://ismir2015.uma.es/LBD/LBD27.pdf Music/Noise segmentation No 2D Cross-entropy 2015 unpublished Musical instrument sound classification with deep convolutional neural network using feature fusion approach Park, Taejin and Lee, Taejin https://arxiv.org/ftp/arxiv/papers/1512/1512.07370.pdf Instrument recognition [UIOWA MIS](http://theremin.music.uiowa.edu/mis.html) 2015 inproceedings Environmental sound classification with convolutional neural networks Piczak, Karol J http://karol.piczak.com/papers/Piczak2015-ESC-ConvNet.pdf 2015 inproceedings Exploring data augmentation for improved singing voice detection with neural networks Schlüter, Jan and Grill, Thomas https://grrrr.org/pub/schlueter-2015-ismir.pdf https://github.com/f0k/ismir2015 SVD Inhouse & [Jamendo](http://www.mathieuramona.com/wp/data/jamendo/) & [RWC](https://staff.aist.go.jp/m.goto/RWC-MDB/) CNN Dropout {5%, 10%, 20%} & Noise {Gaussian sigma={0.05, 0.1, 0.2}} & Pitch shift +-{10, 20, 30, 50} & Time stretch +-{10, 20, 30, 50} & Loudness +-{5dB, 10dB, 20dB} & Frequency filter +-{5dB, 10dB, 20dB} & Mix {10%, 20%, 30%, 50%} & Combined & Test and train Spectrogram 2015 techreport Singer traits identification using deep neural network Shi, Zhengshan https://cs224d.stanford.edu/reports/SkiZhengshan.pdf 2015 inproceedings A hybrid recurrent neural network for music transcription Sigtia, Siddharth and Benetos, Emmanouil and Boulanger-Lewandowski, Nicolas and Weyde, Tillman and Garcez, Artur S d'Avila and Dixon, Simon https://arxiv.org/pdf/1411.1623.pdf Transcription [MAPS](http://www.tsi.telecom-paristech.fr/aao/en/2010/07/08/maps-database-a-piano-database-for-multipitch-estimation-and-automatic-transcription-of-music/) RNN No No No No 2015 unpublished An end-to-end neural network for polyphonic music transcription Sigtia, Siddharth and Benetos, Emmanouil and Dixon, Simon https://arxiv.org/pdf/1508.01774.pdf Transcription CQT 2015 unpublished Deep karaoke: Extracting vocals from musical mixtures using a convolutional deep neural network Simpson, Andrew J. R. and Roma, Gerard and Plumbley, Mark D. https://link.springer.com/chapter/10.1007/978-3-319-22482-4_50 Source separation [MedleyDB](http://medleydb.weebly.com/) STFT 2015 inproceedings Folk music style modelling by recurrent neural networks with long short term memory units Sturm, Bob L. and Santos, João Felipe and Korshunova, Iryna http://ismir2015.uma.es/LBD/LBD13.pdf https://github.com/IraKorshunova/folk-rnn Composition 2015 inproceedings Deep neural network based instrument extraction from music Uhlich, Stefan and Giron, Franck and Mitsufuji, Yuki https://www.researchgate.net/profile/Stefan_Uhlich/publication/282001406_Deep_neural_network_based_instrument_extraction_from_music/links/5600eeda08ae07629e52b397/Deep-neural-network-based-instrument-extraction-from-music.pdf Source separation STFT 2015 inproceedings A deep neural network for modeling music Zhang, Pengjing and Zheng, Xiaoqing and Zhang, Wenqiang and Li, Siyan and Qian, Sheng and He, Wenqi and Zhang, Shangtong and Wang, Ziyuan https://www.researchgate.net/profile/Xiaoqing_Zheng3/publication/275347034_A_Deep_Neural_Network_for_Modeling_Music/links/5539d2060cf2239f4e7dad0d/A-Deep-Neural-Network-for-Modeling-Music.pdf MGR CNN 2016 article An efficient approach for segmentation, feature extraction and classification of audio signals Arumugam, Muthumari and Kaliappan, Mala http://file.scirp.org/pdf/CS_2016042615054817.pdf MGR & Instrument recognition [GTzan](http://marsyas.info/downloads/datasets.html) PNN 2016 inproceedings Text-based LSTM networks for automatic music composition Choi, Keunwoo and Fazekas, György and Sandler, Mark Brian https://drive.google.com/file/d/0B1OooSxEtl0FcG9MYnY2Ylh5c0U/view Composition 2016 unpublished Towards playlist generation algorithms using RNNs trained on within-track transitions Choi, Keunwoo and Fazekas, György and Sandler, Mark Brian https://arxiv.org/pdf/1606.02096.pdf Playlist generation RNN 2016 inproceedings Automatic tagging using deep convolutional neural networks Choi, Keunwoo and Fazekas, György and Sandler, Mark Brian https://arxiv.org/pdf/1606.00298.pdf MGR FCN 2016 inproceedings Automatic chord estimation on seventhsbass chord vocabulary using deep neural network Deng, Junqi and Kwok, Yu-Kwong http://ieeexplore.ieee.org/abstract/document/7471677/ Chord recognition 2016 inproceedings DeepBach: A steerable model for Bach chorales generation Hadjeres, Gaëtan and Pachet, François https://arxiv.org/pdf/1612.01010.pdf https://github.com/Ghadjeres/DeepBach 2016 inproceedings Bayesian meter tracking on learned signal representations Holzapfel, Andre and Grill, Thomas http://www.rhythmos.org/MMILab-Andre_files/ISMIR2016_CNNDBNbeats_camready.pdf Beat detection CNN 2016 unpublished Deep learning for music Huang, Allen and Wu, Raymond https://arxiv.org/pdf/1606.04930.pdf Composition [Bach Corpus](http://musedata.org/) RNN-LSTM 2016 inproceedings Learning temporal features using a deep neural network and its application to music genre classification Jeong, Il-Young and Lee, Kyogu https://www.researchgate.net/profile/Il_Young_Jeong/publication/305683876_Learning_temporal_features_using_a_deep_neural_network_and_its_application_to_music_genre_classification/links/5799a27c08aec89db7bb9f92.pdf STFT & Cepstrum 2016 unpublished On the potential of simple framewise approaches to piano transcription Kelz, Rainer and Dorfer, Matthias and Korzeniowski, Filip and Böck, Sebastian and Arzt, Andreas and Widmer, Gerhard https://arxiv.org/pdf/1612.05153.pdf DNN & ConvNet 2016 inproceedings Feature learning for chord recognition: The deep chroma extractor Korzeniowski, Filip and Widmer, Gerhard https://arxiv.org/pdf/1612.05065.pdf https://github.com/fdlm/chordrec/tree/master/experiments/ismir2016 Chord recognition 2016 inproceedings A fully convolutional deep auditory model for musical chord recognition Korzeniowski, Filip and Widmer, Gerhard https://www.researchgate.net/profile/Filip_Korzeniowski/publication/305590295_A_Fully_Convolutional_Deep_Auditory_Model_for_Musical_Chord_Recognition/links/579486ba08aed51475cc6958/A-Fully-Convolutional-Deep-Auditory-Model-for-Musical-Chord-Recognition.pdf?_iepl%5BhomeFeedViewId%5D=HTzFFmKPia2YminQ4psHT5at&_iepl%5Bcontexts%5D%5B0%5D=pcfhf&_iepl%5BinteractionType%5D=publicationDownload&origin=publication_detail&ev=pub_int_prw_xdl&msrp=Dz_6LKHzYcPyP-LmgZPF-m63ayZ6k0entFEntooiu_e32zfETNQXKPQSTFOI87NONIIQuUQdnUtwORdomTXfteTrb09KiAIdDtBJnw_02P6JeRr5zu2eyaCG.2Uxsi_eENxtbYL39lvorIK8LofRYhkgpUHzpzmVzkIEiyHc0wUY87rEa4PH1qbXi4k4RyagHUsA2IsZtewnprglORjx2v9Cwbk9ZfQ.cd67BaqtHul_hE6SX6vUFKuldz81aH6dWq-cYMkq5vQKCHcvB8l9zgeM694Efb_r2wBB5GT9idt3OLeME0UxVHI6ROxamgK3LMNlSw.JtZXAo9HhR9t-8Wl3gxJgnoM4--rtmDEUDbXSWezbFyU-CoB_nyfxbRQ4kdoN4-5aJ3Tgx4YHdikicqAhc_cezB2ZntjxkB4rEDx1A Chord recognition 2016 inproceedings A deep bidirectional long short-term memory based multi-scale approach for music dynamic emotion prediction Li, Xinxing and Xianyu, Haishu and Tian, Jiashen and Chen, Wenxiao and Meng, Fanhang and Xu, Mingxing and Cai, Lianhong http://ieeexplore.ieee.org/document/7471734/ MER RNN & BILSTM & ELM 2016 inproceedings Event localization in music auto-tagging Liu, Jen-Yu and Yang, Yi-Hsuan http://mac.citi.sinica.edu.tw/~yang/pub/liu16mm.pdf https://github.com/ciaua/clip2frame CNN 2016 inproceedings Deep convolutional networks on the pitch spiral for musical instrument recognition Lostanlen, Vincent and Cella, Carmine-Emanuele https://github.com/lostanlen/ismir2016/blob/master/paper/lostanlen_ismir2016.pdf https://github.com/lostanlen/ismir2016 Instrument recognition 2016 inproceedings SampleRNN: An unconditional end-to-end neural audio generation model Mehri, Soroush and Kumar, Kundan and Gulrajani, Ishaan and Kumar, Rithesh and Jain, Shubham and Sotelo, Jose and Courville, Aaron and Bengio, Yoshua https://openreview.net/pdf?id=SkxKPDv5xl https://github.com/soroushmehr/sampleRNN_ICLR2017 Composition [32 Beethoven’s piano sonatas gathered from https://archive.org](https://soundcloud.com/samplernn/sets) RNN 2016 unpublished Robust audio event recognition with 1-max pooling convolutional neural networks Phan, Huy and Hertel, Lars and Maass, Marco and Mertins, Alfred https://arxiv.org/pdf/1604.06338.pdf Event recognition [RWC](https://staff.aist.go.jp/m.goto/RWC-MDB/) CNN 2016 inproceedings Experimenting with musically motivated convolutional neural networks Pons, Jordi and Lidy, Thomas and Serra, Xavier http://jordipons.me/media/CBMI16.pdf https://github.com/jordipons/ [Ballroom](http://mtg.upf.edu/ismir2004/contest/tempoContest/node5.html) 2016 inproceedings Singing voice melody transcription using deep neural networks Rigaud, François and Radenen, Mathieu https://wp.nyu.edu/ismir2016/wp-content/uploads/sites/2294/2016/07/163_Paper.pdf F0 & VAD DNN & RNN-LSTM 2016 inproceedings Singing voice separation using deep neural networks and F0 estimation Roma, Gerard and Grais, Emad M. and Simpson, Andrew J. R. and Plumbley, Mark D. http://www.music-ir.org/mirex/abstracts/2016/RSGP1.pdf http://cvssp.org/projects/maruss/mirex2016/ SVS [iKala](http://mac.citi.sinica.edu.tw/ikala/) 2016 inproceedings Learning to pinpoint singing voice from weakly labeled examples Schlüter, Jan http://www.ofai.at/~jan.schlueter/pubs/2016_ismir.pdf CNN STFT 2016 inproceedings Analysis of time-frequency representations for musical onset detection with convolutional neural network Stasiak, Bartłomiej and Mońko, Jędrzej http://ieeexplore.ieee.org/abstract/document/7733228/ Onset detection 2016 article Note onset detection in musical signals via neural-network-based multi-ODF fusion Stasiak, Bartłomiej and Mońko, Jędrzej and Niewiadomski, Adam https://www.degruyter.com/downloadpdf/j/amcs.2016.26.issue-1/amcs-2016-0014/amcs-2016-0014.pdf Onset detection Inhouse & [RWC](https://staff.aist.go.jp/m.goto/RWC-MDB/) NNMODFF Onset activation 2016 inproceedings Music transcription modelling and composition using deep learning Sturm, Bob L. and Santos, João Felipe and Ben-Tal, Oded and Korshunova, Iryna https://drive.google.com/file/d/0B1OooSxEtl0FcTBiOGdvSTBmWnc/view https://github.com/IraKorshunova/folk-rnn Composition 2016 inproceedings Convolutional neural network for robust pitch determination Su, Hong and Zhang, Hui and Zhang, Xueliang and Gao, Guanglai http://www.mirlab.org/conference_papers/International_Conference/ICASSP%202016/pdfs/0000579.pdf Pitch determination CNN & MLP 2016 unpublished Deep convolutional neural networks and data augmentation for acoustic event detection Takahashi, Naoya and Gygli, Michael and Pfister, Beat and Van Gool, Luc https://arxiv.org/pdf/1604.07160.pdf https://bitbucket.org/naoya1/aenet_release Event recognition [Acoustic Event](https://data.vision.ee.ethz.ch/cvl/ae_dataset/) CNN Mixing 2017 unpublished Gabor frames and deep scattering networks in audio processing Bammer, Roswitha and Doerfler, Monika https://arxiv.org/pdf/1706.08818.pdf 2017 inproceedings Vision-based detection of acoustic timed events: A case study on clarinet note onsets Bazzica, Alessio and Van Gemert, JC and Liem, CCS and Hanjalic, A http://dorienherremans.com/dlm2017/papers/bazzica2017clarinet.pdf Onset detection [C4S](http://mmc.tudelft.nl/users/alessio-bazzica#C4S-dataset) CNN 2017 unpublished Deep learning techniques for music generation - A survey Briot, Jean-Pierre and Hadjeres, Gaëtan and Pachet, François https://arxiv.org/pdf/1709.01620.pdf Survey & Composition [JSB Chorales](https://github.com/czhuang/JSB-Chorales-dataset) & [MusicNet](https://homes.cs.washington.edu/~thickstn/musicnet.html) & [Symbolic music data](http://users.cecs.anu.edu.au/~christian.walder/) & [LSDB](lsdb.flow-machines.com/) No 2017 inproceedings JamBot: Music theory aware chord based generation of polyphonic music with LSTMs Brunner, Gino and Wang, Yuyi and Wattenhofer, Roger and Wiesendanger, Jonas https://arxiv.org/pdf/1711.07682.pdf https://github.com/brunnergino/JamBot Composition [Lakh MIDI](https://labrosa.ee.columbia.edu/sounds/music/) Keras-TensorFlow RNN-LSTM No No 4 No Softmax 0.00001 Adam 1 2017 unpublished XFlow: 1D <-> 2D cross-modal deep neural networks for audiovisual classification Cangea, Cătălina and Veličković, Petar and Liò, Pietro https://arxiv.org/pdf/1709.00572.pdf 2017 inproceedings Machine listening intelligence Cella, Carmine-Emanuele http://dorienherremans.com/dlm2017/papers/cella2017mli.pdf No Manifesto No No No No No No No No No No No No 2017 inproceedings Monoaural audio source separation using deep convolutional neural networks Chandna, Pritish and Miron, Marius and Janer, Jordi and Gómez, Emilia http://mtg.upf.edu/system/files/publications/monoaural-audio-source_0.pdf https://github.com/MTG/DeepConvSep Source separation [DSD100](http://sisec17.audiolabs-erlangen.de/#/dataset) CNN 2017 inproceedings Deep multimodal network for multi-label classification Chen, Tanfang and Wang, Shangfei and Chen, Shiyu http://ieeexplore.ieee.org/abstract/document/8019322/ General audio classification 2017 unpublished A tutorial on deep learning for music information retrieval Choi, Keunwoo and Fazekas, György and Cho, Kyunghyun and Sandler, Mark Brian https://arxiv.org/pdf/1709.04396.pdf https://github.com/keunwoochoi/dl4mir General audio classification 2017 unpublished A comparison on audio signal preprocessing methods for deep neural networks on music tagging Choi, Keunwoo and Fazekas, György and Cho, Kyunghyun and Sandler, Mark Brian https://arxiv.org/pdf/1709.01922.pdf https://github.com/keunwoochoi/transfer_learning_music MGR [MSD](https://labrosa.ee.columbia.edu/millionsong/) 2017 inproceedings Transfer learning for music classification and regression tasks Choi, Keunwoo and Fazekas, György and Sandler, Mark Brian and Cho, Kyunghyun https://arxiv.org/pdf/1703.09179v3.pdf https://github.com/keunwoochoi/transfer_learning_music General audio classification [MSD](https://labrosa.ee.columbia.edu/millionsong/) No Adam 2017 inproceedings Convolutional recurrent neural networks for music classification Choi, Keunwoo and Fazekas, György and Sandler, Mark Brian and Cho, Kyunghyun http://ieeexplore.ieee.org/abstract/document/7952585/ https://github.com/keunwoochoi/icassp_2017 MGR Models & split sets only CRNN 2017 article An evaluation of convolutional neural networks for music classification using spectrograms Costa, Yandre MG and Oliveira, Luiz S and Silla, Carlos N http://www.inf.ufpr.br/lesoliveira/download/ASOC2017.pdf General audio classification [LMD](https://sites.google.com/site/carlossillajr/resources/the-latin-music-database-lmd) Caffe CNN 128 0.001 Tesla C2050 2017 unpublished Large vocabulary automatic chord estimation using deep neural nets: Design framework, system variations and limitations Deng, Junqi and Kwok, Yu-Kwong https://arxiv.org/pdf/1709.07153.pdf Chord recognition 2017 unpublished Basic filters for convolutional neural networks: Training or design? Doerfler, Monika and Grill, Thomas and Bammer, Roswitha and Flexer, Arthur https://arxiv.org/pdf/1709.02291.pdf SVD Inhouse Raw & Mel-spectrogram 0.001 Adam 2017 unpublished Ensemble Of Deep Neural Networks For Acoustic Scene Classification Duppada, Venkatesh and Hiray, Sushant https://arxiv.org/pdf/1708.05826.pdf 2017 article Robust downbeat tracking using an ensemble of convolutional networks Durand, Simon and Bello, Juan Pablo and David, Bertrand and Richard, Gaël http://ieeexplore.ieee.org/abstract/document/7728057/ Beat detection CNN 2017 inproceedings Music signal processing using vector product neural networks Fan, Zhe-Cheng and Chan, TS and Yang, Yi-Hsuan and Jang, Jyh-Shing R http://dorienherremans.com/dlm2017/papers/fan2017vector.pdf SVS [iKala](http://mac.citi.sinica.edu.tw/ikala/) VPNN & DNN 2017 inproceedings Transforming musical signals through a genre classifying convolutional neural network Geng, Shijia and Ren, Gang and Ogihara, Mitsunori http://dorienherremans.com/dlm2017/papers/geng2017genre.pdf Composition Inhouse CNN 2017 inproceedings Audio to score matching by combining phonetic and duration information Gong, Rong and Pons, Jordi and Serra, Xavier https://arxiv.org/pdf/1707.03547.pdf https://github.com/ronggong/jingjuSingingPhraseMatching/tree/v0.1.0 2017 unpublished Interactive music generation with positional constraints using anticipation-RNNs Hadjeres, Gaëtan and Nielsen, Frank https://arxiv.org/pdf/1709.06404.pdf Composition ARNN 2017 unpublished Deep rank-based transposition-invariant distances on musical sequences Hadjeres, Gaëtan and Nielsen, Frank https://arxiv.org/pdf/1709.00740.pdf 2017 unpublished GLSR-VAE: Geodesic latent space regularization for variational autoencoder architectures Hadjeres, Gaëtan and Nielsen, Frank and Pachet, François https://arxiv.org/pdf/1707.04588.pdf 2017 article Deep convolutional neural networks for predominant instrument recognition in polyphonic music Han, Yoonchang and Kim, Jaehun and Lee, Kyogu and Han, Yoonchang and Kim, Jaehun and Lee, Kyogu http://dl.acm.org/citation.cfm?id=3068697 Instrument recognition [IRMAS](https://www.upf.edu/web/mtg/irmas) CNN 2017 inproceedings CNN architectures for large-scale audio classification Hershey, Shawn and Chaudhuri, Sourish and Ellis, Daniel P. W. and Gemmeke, Jort F. and Jansen, Aren and Moore, R. Channing and Plakal, Manoj and Platt, Devin and Saurous, Rif A. and Seybold, Bryan and Slaney, Malcolm and Weiss, Ron J. and Wilson, Kevin https://arxiv.org/pdf/1609.09430v2.pdf CNN 2017 inproceedings DeepSheet: A sheet music generator based on deep learning Hsu, Yu-Lun and Lin, Chi-Po and Lin, Bo-Chen and Kuo, Hsu-Chan and Cheng, Wen-Huang and Hu, Min-Chun http://ieeexplore.ieee.org/abstract/document/8026272/ Composition 2017 inproceedings Talking Drums: Generating drum grooves with neural networks Hutchings, P. http://dorienherremans.com/dlm2017/papers/hutchings2017drums.pdf Composition RNN 2017 inproceedings Singing voice separation with deep U-Net convolutional networks Jansson, Andreas and Humphrey, Eric J. and Montecchio, Nicola and Bittner, Rachel and Kumar, Aparna and Weyde, Tillman https://ismir2017.smcnus.org/wp-content/uploads/2017/10/171_Paper.pdf https://github.com/Xiao-Ming/UNet-VocalSeparation-Chainer SVS No [iKala](http://mac.citi.sinica.edu.tw/ikala/) & [MedleyDB](http://medleydb.weebly.com/) CNN & U-Net No Sigmoid Adam 2017 inproceedings Music emotion recognition via end-to-end multimodal neural networks Jeon, Byungsoo and Kim, Chanju and Kim, Adrian and Kim, Dongwon and Park, Jangyeon and Ha, Jung-Woo http://ceur-ws.org/Vol-1905/recsys2017_poster18.pdf MER Inhouse CNN 2017 inproceedings Chord label personalization through deep learning of integrated harmonic interval-based representations Koops, Hendrik Vincent and De Haas, W Bas and Bransen, Jeroen and Volk, Anja http://dorienherremans.com/dlm2017/papers/koops2017pers.pdf 2017 unpublished End-to-end musical key estimation using a convolutional neural network Korzeniowski, Filip and Widmer, Gerhard https://arxiv.org/pdf/1706.02921.pdf 2017 inproceedings MediaEval 2017 AcousticBrainz genre task: Multilayer perceptron approach Koutini, Khaled and Imenina, Alina and Dorfer, Matthias and Gruber, Alexander Rudolf and Schedl, Markus http://www.cp.jku.at/research/papers/Koutini_2017_mediaeval-acousticbrainz.pdf 2017 unpublished Classification-based singing melody extraction using deep convolutional neural networks Kum, Sangeun and Nam, Juhan https://www.preprints.org/manuscript/201711.0027/v1 No F0 No [LabROSA](http://labrosa.ee.columbia.edu/projects/melody/) & [MedleyDB](http://medleydb.weebly.com/) & [Jamendo](http://www.mathieuramona.com/wp/data/jamendo/) & [RWC](https://staff.aist.go.jp/m.goto/RWC-MDB/) & [iKala](http://mac.citi.sinica.edu.tw/ikala/) & [MIR-1K](https://sites.google.com/site/unvoicedsoundseparation/mir-1k) & [ADC2004](http://labrosa.ee.columbia.edu/projects/melody/) Keras CNN 0.3 No 100 Pitch shift -2, -1, +1, +2 semitones Leaky ReLU 0.02 SGD 2 2017 article Multi-level and multi-scale feature aggregation using pre-trained convolutional neural networks for music auto-tagging Lee, Jongpil and Nam, Juhan https://arxiv.org/pdf/1703.01793v2.pdf 2017 unpublished Multi-level and multi-scale feature aggregation using sample-level deep convolutional neural networks for music classification Lee, Jongpil and Nam, Juhan https://arxiv.org/pdf/1706.06810.pdf https://github.com/jongpillee/musicTagging_MSD [MSD](https://labrosa.ee.columbia.edu/millionsong/) 2017 inproceedings Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms Lee, Jongpil and Park, Jiyoung and Kim, Keunhyoung Luke and Nam, Juhan https://arxiv.org/pdf/1703.01789v2.pdf 2017 unpublished A SeqGAN for Polyphonic Music Generation Lee, Sang-gil and Hwang, Uiwon and Min, Seonwoo and Yoon, Sungroh https://arxiv.org/pdf/1710.11418.pdf https://github.com/L0SG/seqgan-music Polyphonic music sequence modelling No [Nottingham dataset](http://abc.sourceforge.net/NMD/) Tensorflow SeqGAN No No 100 No MIDI 1D No No 0.01 & 0.001 & 0.0001 No No 2017 inproceedings Harmonic and percussive source separation using a convolutional auto encoder Lim, Wootaek and Lee, Taejin http://www.eurasip.org/Proceedings/Eusipco/Eusipco2017/papers/1570346835.pdf Source separation [DSD100](http://sisec17.audiolabs-erlangen.de/#/dataset) CNN 2017 unpublished Stacked convolutional and recurrent neural networks for music emotion recognition Malik, Miroslav and Adavanne, Sharath and Drossos, Konstantinos and Virtanen, Tuomas and Ticha, Dasa and Jarina, Roman https://arxiv.org/pdf/1706.02292.pdf No MER [Free music archive](http://freemusicarchive.org/) & [MedleyDB](http://medleydb.weebly.com/) & [Jamendo](http://www.mathieuramona.com/wp/data/jamendo/) Keras-Theano CRNN No RMSE Adam 2017 mastersthesis A deep learning approach to source separation and remixing of hiphop music Martel Baro, Héctor https://repositori.upf.edu/bitstream/handle/10230/32919/Martel_2017.pdf?sequence=1&isAllowed=y Source separation & Remixing [DSD100](http://sisec17.audiolabs-erlangen.de/#/dataset) & [HHDS](https://drive.google.com/drive/folders/0B1zpiGdDzFNlbmJyYU1VVFR3OEE) DNN & CNN & RNN Mixing & Circular Shift & Instrument augmentation 2017 inproceedings Music Genre Classification Using Masked Conditional Neural Networks Medhat, Fady and Chesmore, David and Robinson, John https://link.springer.com/chapter/10.1007%2F978-3-319-70096-0_49 No MGR [Ballroom](http://mtg.upf.edu/ismir2004/contest/tempoContest/node5.html) & [Homburg](http://www-ai.cs.uni-dortmund.de/audio.html) Not disclosed MCLNN & CLNN No No No No No Adam No 2017 unpublished Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency Mask Mimilakis, Stylianos Ioannis and Drossos, Konstantinos and Santos, João Felipe and Schuller, Gerald and Virtanen, Tuomas and Bengio, Yoshua https://arxiv.org/pdf/1711.01437.pdf https://github.com/Js-Mim/mss_pytorch SVS [MedleyDB](http://medleydb.weebly.com/) & [DSD100](http://sisec17.audiolabs-erlangen.de/#/dataset) PyTorch RNN & DNN 16 100 No ReLU 0.0001 Adam No 2017 inproceedings Generating data to train convolutional neural networks for classical music source separation Miron, Marius and Janer Mestres, Jordi and Gómez Gutiérrez, Emilia https://www.researchgate.net/profile/Marius_Miron/publication/318322107_Generating_data_to_train_convolutional_neural_networks_for_classical_music_source_separation/links/59637cc3458515a3575b93c6/Generating-data-to-train-convolutional-neural-networks-for-classical-music-source-separation.pdf?_iepl%5BhomeFeedViewId%5D=WchoMnlUL1Hk9hBLVTeR8Amh&_iepl%5Bcontexts%5D%5B0%5D=pcfhf&_iepl%5BinteractionType%5D=publicationDownload&origin=publication_detail&ev=pub_int_prw_xdl&msrp=p3lQ8M4uZlb4TF5Hv9a2U3P2y4wW7ant5KWj4E5-OcD1Mg53p1ykTKHMG9_zVTB9n6mI8fvZOCL2Xhpru186pCEY-2ZxiYR-CB8_QvwHc1kUG-QE4SHdProR.LoJb2BDOiiQth3iR9xgZUxxCWEJgtTBF4whFrFa01OD49-3YYRxA0WQVN--zhtQU_7C2Pt0rKdwoFxT1pfxFvnKXSXmy2eT1Jpz-pw.U1QLoFO_Uc6aQVr2Nm2FcAi6BqAUfngH2Or5__6wegbCgVvTYoIGt22tmCkYbGTOQ_4PxBgt1LrvsFQiL0oMyogP8Yk8myTj0gs9jw.fGpkufGqAI4R2v8Hfe0ThcXL7M7yN2PuAlx974BGVn50SdUWvNhhIPWBD-zWTn8NKtVJx3XrjKXFrMgi9Cx7qGrNP8tBWpha6Srf6g https://github.com/MTG/DeepConvSep Source separation [RWC](https://staff.aist.go.jp/m.goto/RWC-MDB/) & [Bach10](http://music.cs.northwestern.edu/data/Bach10.html) CNN 32 2017 inproceedings Monaural score-informed source separation for classical music using convolutional neural networks Miron, Marius and Janer, Jordi and Gómez, Emilia https://www.researchgate.net/profile/Marius_Miron/publication/318637038_Monaural_score-informed_source_separation_for_classical_music_using_convolutional_neural_networks/links/597327c6458515e26dfdb007/Monaural-score-informed-source-separation-for-classical-music-using-convolutional-neural-networks.pdf?_iepl%5BhomeFeedViewId%5D=WchoMnlUL1Hk9hBLVTeR8Amh&_iepl%5Bcontexts%5D%5B0%5D=pcfhf&_iepl%5BinteractionType%5D=publicationDownload&origin=publication_detail&ev=pub_int_prw_xdl&msrp=Hp6dDqMepEiRZ5E6WkreaqyjFkFkwMxPFoJvr14etVJsoKZBc5qb99fBnJjVUZrRHLFRhaXvNY9k1sMvYPOouuGbQP0YhEGm28zLw_55Zewu86WGnHck1Tqi.93HH2WqXfTedn6IaZRjjhQGYZVDHBz1X6nr4ABBgMAVv584gvGN3sW5IyBAY-4MBWf5DJFPBGm8zsaC2dKz8G-odZPfosWoXY0afAQ.KoCP2mO9l31lCER0oMZMZBrbuRGvb6ZzeBwHb88pL8AhMfJk03Hj1eLrohQIjPDETBj4hhqb0gniDGJgtZ9GnW64ZNjh9GbQDrIl5A.egNQTyC7t8P26zCQWrbEhf51Pxy2JRBZoTkH6SpRHHhRhFl1_AT_AT481lMcFI34-JbeRq-5oTQR7DpvAuw7iUIivd78ltuxpI9syg https://github.com/MTG/DeepConvSep Source separation [Bach10](http://music.cs.northwestern.edu/data/Bach10.html) CNN 2017 inproceedings Multi-label music genre classification from audio, text, and images using deep features Oramas, Sergio and Nieto, Oriol and Barbieri, Francesco and Serra, Xavier https://ismir2017.smcnus.org/wp-content/uploads/2017/10/126_Paper.pdf https://github.com/sergiooramas/tartarus MGR [MSD](https://labrosa.ee.columbia.edu/millionsong/) CNN No 2017 inproceedings A deep multimodal approach for cold-start music recommendation Oramas, Sergio and Nieto, Oriol and Sordo, Mohamed and Serra, Xavier https://arxiv.org/pdf/1706.09739.pdf https://github.com/sergiooramas/tartarus Recommendation [MSD](https://labrosa.ee.columbia.edu/millionsong/) CNN 32 2017 inproceedings Melody extraction and detection through LSTM-RNN with harmonic sum loss Park, Hyunsin and Yoo, Chang D http://ieeexplore.ieee.org/abstract/document/7952660/ Artist recognition & MGR No 2017 unpublished Representation learning of music using artist labels Park, Jiyoung and Lee, Jongpil and Park, Jangyeon and Ha, Jung-Woo and Nam, Juhan https://arxiv.org/pdf/1710.06648.pdf [MSD](https://labrosa.ee.columbia.edu/millionsong/) & [GTzan](http://marsyas.info/downloads/datasets.html) & [Magnatagatune](http://mirg.city.ac.uk/codeapps/the-magnatagatune-dataset) CNN 2017 inproceedings Toward inverse control of physics-based sound synthesis Pfalz, A and Berdahl, E http://dorienherremans.com/dlm2017/papers/pfalz2017synthesis.pdf https://www.cct.lsu.edu/~apfalz/inverse_control.html 2017 techreport DNN and CNN with weighted and multi-task loss functions for audio event detection Phan, Huy and Krawczyk-Becker, Martin and Gerkmann, Timo and Mertins, Alfred https://arxiv.org/pdf/1708.03211.pdf Event recognition CNN & DNN 2017 inproceedings Score-informed syllable segmentation for a cappella singing voice with convolutional neural networks Pons, Jordi and Gong, Rong and Serra, Xavier https://ismir2017.smcnus.org/wp-content/uploads/2017/10/46_Paper.pdf https://github.com/ronggong/jingjuSyllabicSegmentaion/tree/v0.1.0 Syllable segmentation 128 Adam 2017 inproceedings End-to-end learning for music audio tagging at scale Pons, Jordi and Nieto, Oriol and Prockup, Matthew and Schmidt, Erik M. and Ehmann, Andreas F. and Serra, Xavier https://arxiv.org/pdf/1711.02520.pdf https://github.com/jordipons/music-audio-tagging-at-scale-models General audio classification Inhouse Tensorflow CNN 16 No No ReLU 0.001 Adam No 2017 inproceedings Designing efficient architectures for modeling temporal features with convolutional neural networks Pons, Jordi and Serra, Xavier http://ieeexplore.ieee.org/document/7952601/ https://github.com/jordipons/ICASSP2017 MGR [Ballroom](http://mtg.upf.edu/ismir2004/contest/tempoContest/node5.html) CNN 2017 inproceedings Timbre analysis of music audio signals with convolutional neural networks Pons, Jordi and Slizovskaia, Olga and Gong, Rong and Gómez, Emilia and Serra, Xavier https://github.com/ronggong/EUSIPCO2017 https://github.com/jordipons/EUSIPCO2017 CNN 2017 inproceedings The MUSDB18 corpus for music separation Rafii, Zafar and Liutkus, Antoine and Stöter, Fabian-Robert and Mimilakis, Stylianos Ioannis and Bittner, Rachel https://doi.org/10.5281/zenodo.1117372 https://github.com/sigsep/website Source separation [MUSDB18](https://sigsep.github.io/datasets/musdb.html) No No 2017 inproceedings Deep learning and intelligent audio mixing Ramírez, Marco A. Martínez and Reiss, Joshua D. http://www.semanticaudio.co.uk/wp-content/uploads/2017/09/WIMP2017_Martinez-RamirezReiss.pdf No Mixing [Open Multitrack Testbed](http://www.semanticaudio.co.uk/projects/omtb/) DAE No Adam 2017 phdthesis Deep learning for event detection, sequence labelling and similarity estimation in music signals Schlüter, Jan http://ofai.at/~jan.schlueter/pubs/phd/phd.pdf 2017 inproceedings Music feature maps with convolutional neural networks for music genre classification Senac, Christine and Pellegrini, Thomas and Mouret, Florian and Pinquier, Julien https://www.researchgate.net/profile/Thomas_Pellegrini/publication/319326354_Music_Feature_Maps_with_Convolutional_Neural_Networks_for_Music_Genre_Classification/links/59ba5ae3458515bb9c4c6724/Music-Feature-Maps-with-Convolutional-Neural-Networks-for-Music-Genre-Classification.pdf?origin=publication_detail&ev=pub_int_prw_xdl&msrp=wzXuHZAa5zAnqEmErYyZwIRr2H0q01LnNEd4Wd7A15CQfdVLwdy98pmE-AdnrDvoc3-bVENSFrHt0yhaOiE2mQrYllVS9CJZOk-c9R0j_R1rbgcZugS6RtQ_.AUjPuJSF5P_DMngf-woH7W-7jdnQlbNQziR4_h6NnCHfR_zGcEa8vOyyOz5gx5nc4azqKTPQ5ZgGGLUxkLj1qCQLEQ5ThkhGlWHLyA.s6MBZE20-EO_RjRGCOCV4wk0WSFdN56Aloiraxz9hKCbJwRM2Et27RHVUA8jj9H8qvXIB6f7zSIrQgjXGrL2yCpyQlLffuf57rzSwg.KMMXbZrHsihV8DJM53xkHAWf3VebCJESi4KU4btNv9nQsyK2KnkhSQaTILKv0DSZY3c70a61LzywCBuoHtIhVOFhW5hVZN2n5O9uKQ MGR [GTzan](http://marsyas.info/downloads/datasets.html) CNN Spectrograms & common audio features 2017 inproceedings Automatic drum transcription for polyphonic recordings using soft attention mechanisms and convolutional neural networks Southall, Carl and Stables, Ryan and Hockman, Jason https://carlsouthall.files.wordpress.com/2017/12/ismir2017adt.pdf https://github.com/CarlSouthall/ADTLib Transcription [IDMT-SMT-Drums](https://www.idmt.fraunhofer.de/en/business_units/m2d/smt/drums.html) CNN & BRNN 2017 unpublished Adversarial semi-supervised audio source separation applied to singing voice extraction Stoller, Daniel and Ewert, Sebastian and Dixon, Simon https://arxiv.org/pdf/1711.00048.pdf No Source separation No [iKala](http://mac.citi.sinica.edu.tw/ikala/) & [MedleyDB](http://medleydb.weebly.com/) & [DSD100](http://sisec17.audiolabs-erlangen.de/#/dataset) & [CCMixter](https://members.loria.fr/ALiutkus/kam/) CNN & U-Net No 6 No ReLU & Leaky ReLU Adam NVIDIA GTX1080 2017 article Taking the models back to music practice: Evaluating generative transcription models built using deep learning Sturm, Bob L. and Ben-Tal, Oded http://jcms.org.uk/issues/Vol2Issue1/taking-models-back-to-music-practice/Taking%20the%20Models%20back%20to%20Music%20Practice:%20Evaluating%20Generative%20Transcription%20Models%20built%20using%20Deep%20Learning.pdf https://github.com/IraKorshunova/folk-rnn Composition 2017 inproceedings Generating nontrivial melodies for music as a service Teng, Yifei and Zhao, An and Goudeseune, Camille https://ismir2017.smcnus.org/wp-content/uploads/2017/10/178_Paper.pdf Composition 2017 unpublished Invariances and data augmentation for supervised music transcription Thickstun, John and Harchaoui, Zaid and Foster, Dean and Kakade, Sham M. https://arxiv.org/pdf/1711.04845.pdf https://github.com/jthickstun/thickstun2018invariances/ Transcription [MusicNet](https://homes.cs.washington.edu/~thickstn/musicnet.html) Tensorflow CNN 150 No Pitch shift integer in [-5, 5] semitones and continuous in [-0.1, 0.1] No No No 1 NVIDIA 1080Ti 2017 inproceedings Lyrics-based music genre classification using a hierarchical attention network Tsaptsinos, Alexandros https://ismir2017.smcnus.org/wp-content/uploads/2017/10/43_Paper.pdf https://github.com/alexTsaptsinos/lyricsHAN MGR [LyricFind](http://lyricfind.com/) Tensorflow HAN cross-entropy 2017 unpublished A hybrid DSP/deep learning approach to real-time full-band speech enhancement Valin, Jean-Marc https://arxiv.org/pdf/1709.08243.pdf https://github.com/xiph/rnnoise/ Noise suppression [TSP](http://www-mmsp.ece.mcgill.ca/Documents/Data/) & [NTT MLS](http://www.ntt-at.com/product/speech/) RNN No BFCC (22), 1st and 2nd derivatives of first 6 BFCCs, 6 coefficients of DCT of pitch correlation, pitch period, spectral non-stationary metric Custom 2017 phdthesis Convolutional methods for music analysis Velarde, Gissel http://vbn.aau.dk/files/260308151/PHD_Gissel_Velarde_E_pdf.pdf 2017 inproceedings Extending temporal feature integration for semantic audio analysis Vrysis, Lazaros and Tsipas, Nikolaos and Dimoulas, Charalampos and Papanikolaou, George http://www.aes.org/e-lib/browse.cfm?elib=18682 ANN 2017 inproceedings Recognition and retrieval of sound events using sparse coding convolutional neural network Wang, Chien-Yao and Santoso, Andri and Mathulaprangsan, Seksan and Chiang, Chin-Chin and Wu, Chung-Hsien and Wang, Jia-Ching http://ieeexplore.ieee.org/abstract/document/8019552/ Event recognition CNN 2017 article A two-stage approach to note-level transcription of a specific piano Wang, Qi and Zhou, Ruohua and Yan, Yonghong http://www.mdpi.com/2076-3417/7/9/901/htm Transcription 2017 unpublished Reducing model complexity for DNN based large-scale audio classification Wu, Yuzhong and Lee, Tan https://arxiv.org/pdf/1711.00229.pdf No General audio classification No [AudioSet](https://research.google.com/audioset/index.html) & [TUT Acoustic Scenes 2016](http://www.cs.tut.fi/~mesaros/pubs/mesaros_eusipco2016-dcase.pdf) CNN & RNN & MLP & AlexNet & ResNet No ReLU Adam 2017 inproceedings Audio spectrogram representations for processing with convolutional neural networks Wyse, Lonce http://dorienherremans.com/dlm2017/papers/wyse2017spect.pdf http://lonce.org/research/audioST/ Review & Comparison CNN 2017 article Unsupervised feature learning based on deep models for environmental audio tagging Xu, Yong and Huang, Qiang and Wang, Wenwu and Foster, Peter and Sigtia, Siddharth and Jackson, Philip J. B. and Plumbley, Mark D. https://arxiv.org/pdf/1607.03681.pdf 2017 inproceedings Attention and localization based on a deep convolutional recurrent model for weakly supervised audio tagging Xu, Yong and Kong, Qiuqiang and Huang, Qiang and Wang, Wenwu and Plumbley, Mark D. https://arxiv.org/pdf/1703.06052.pdf https://github.com/yongxuUSTC/att_loc_cgrnn DCASE 2016 Task 4 Domestic audio tagging CRNN 2017 techreport Surrey-CVSSP system for DCASE2017 challenge task4 Xu, Yong and Kong, Qiuqiang and Wang, Wenwu and Plumbley, Mark D. https://www.cs.tut.fi/sgn/arg/dcase2017/documents/challenge_technical_reports/DCASE2017_Xu_146.pdf https://github.com/yongxuUSTC/dcase2017_task4_cvssp Event recognition 2017 inproceedings A study on LSTM networks for polyphonic music sequence modelling Ycart, Adrien and Benetos, Emmanouil https://qmro.qmul.ac.uk/xmlui/handle/123456789/24946 http://www.eecs.qmul.ac.uk/~ay304/code/ismir17 Polyphonic music sequence modelling Inhouse & [Piano-midi.de](Piano-midi.de) RNN-LSTM Pitch shift 2018 inproceedings MuseGAN: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment Dong, Hao-Wen and Hsiao, Wen-Yi and Yang, Li-Chia and Yang, Yi-Hsuan https://arxiv.org/pdf/1709.06298.pdf https://github.com/salu133445/musegan Composition No [Lakh Pianoroll Datase](https://github.com/salu133445/musegan/blob/master/docs/dataset.md) No GAN & CNN No No No No Piano-roll 1D ReLU & Leaky ReLU No No Adam 1 Tesla K40m 2018 unpublished Music transformer: Generating music with long-term structure Huang, Cheng-Zhi Anna and Vaswani, Ashish and Uszkoreit, Jakob and Shazeer, Noam and Simon, Ian and Hawthorne, Curtis and Dai, Andrew M. and Hoffman, Matthew D. and Dinculescu, Monica and Eck, Douglas https://arxiv.org/pdf/1809.04281.pdf Polyphonic music sequence modelling No [J.S. Bach chorales dataset](https://github.com/czhuang/JSB-Chorales-dataset) & [Piano-e-Competition dataset (competition history)](http://www.piano-e-competition.com/) No Transformer & RNN & tensor2tensor 0.1 1 No Time Stretches & pitch transcription MIDI 1D No No 0.1 No No 2018 inproceedings Music theory inspired policy gradient method for piano music transcription Li, Juncheng and Qu, Shuhui and Wang, Yun and Li, Xinjian and Das, Samarjit and Metze, Florian https://nips2018creativity.github.io/doc/music_theory_inspired_policy_gradient.pdf No Transcription No [MAPS](http://www.tsi.telecom-paristech.fr/aao/en/2010/07/08/maps-database-a-piano-database-for-multipitch-estimation-and-automatic-transcription-of-music/) CNN & RNN No 8 No No Log Mel-spectrogram with 48 bins per octave and 512 hop-size and 2018 window size and 16 kHz sample rate 2D No binary cross-entropy 0.0006 Adam No 2019 inproceedings Enabling factorized piano music modeling and generation with the MAESTRO dataset Hawthorne, Curtis and Stasyuk, Andriy and Roberts, Adam and Simon, Ian and Huang, Cheng-Zhi Anna and Dieleman, Sander and Elsen, Erich and Engel, Jesse and Eck, Douglas https://arxiv.org/abs/1810.12247 https://github.com/magenta/magenta/tree/master/magenta/models/onsets_frames_transcription Transcription No [MAESTRO](https://magenta.tensorflow.org/datasets/maestro/) No No No No No pitch-shift {+-0.1 semitones} & compression {0 - 100} & EQ {32 - 4096} & Reverb & {0 - 70} & Pink-noise {0 - 0.04} MIDI 1D No No No No No 2019 unpublished Generating Long Sequences with Sparse Transformers Rewon Child and Scott Gray and Alec Radford and Ilya Sutskever https://arxiv.org/pdf/1904.10509.pdf https://github.com/openai/sparse_attention Audio generation No [413 hours of recorded solo piano music](http://papers.nips.cc/paper/8023-the-challenge-of-realistic-music-generation-modelling-raw-audio-at-scale-supplemental.zip) Tensorflow Transformer 0.25 No 120 No Raw Audio 1D No No 0.00035 Adam 8 NVIDIA Tesla V100 2021 inproceedings DadaGP: a Dataset of Tokenized GuitarPro Songs for Sequence Models Sarmento, Pedro and Kumar, Adarsh and Carr, CJ and Zukowski, Zack and Barthet, Mathieu and Yang, Yi-Hsuan https://archives.ismir.net/ismir2021/paper/000076.pdf https://github.com/dada-bots/dadaGP Polyphonic music sequence modelling [DadaGP](https://drive.google.com/drive/folders/1USNH8olG9uy6vodslM3iXInBT725zult?usp=sharing) No No