Acoustic-based LEGO recognition using attention-based convolutional neural networks
This research presents LegoNet, a compact CNN model that classifies LEGO pieces using their impact sounds when dropped. It achieves over 99% accuracy and is robust even under noise. Using full sound sequences improves recognition, and posture-specific data enhances performance with fewer samples. The system is efficient and suitable for real-world deployment. It opens new possibilities in audio-based object recognition.

Table 11 Results of experiments on the specific-posture dataset (i.e., LEGO-III)
Technology Overview
This study introduces an acoustic-based object recognition approach using free-fall impact sounds of LEGO components. By dropping over 200 LEGO objects from a fixed height onto a plane, the study collects audio data representing impact events. A compact 2D convolutional neural network (CNN) named LegoNet is proposed, featuring frame-level attention and time-distributed fully connected layers. LegoNet classifies audio spectrograms with over 99% accuracy, outperforming larger baseline models in both clean and noisy environments. The study also explores the effect of using full sound sequences versus single impact events and the influence of object posture and data size on model performance.
Applications & Benefits
The proposed method provides a novel, efficient alternative to vision-based object recognition, especially useful in low-light or occluded environments. LegoNet is small, robust, and suitable for real-time deployment in intelligent agents, robots, or sorting systems. Using transfer learning, the model achieves high accuracy even with limited data, making it cost-effective for new object types. Furthermore, its robustness to noise and variations in object posture enhances its usability in real-world industrial or educational applications. This approach offers a new dimension in human–machine interaction by leveraging audio signals for physical environment understanding.
Abstract:
This work investigates the classification of LEGO types using deep learning-based audio classification approaches. The motivation for this investigation is based on the following assumption. If objects of the same shape fall freely from a certain height and hit a fixed plane, the impact sounds will be very similar, so we can distinguish the same types of objects from the others. Applying this idea to LEGO recognition, we collect impact sounds of 200 LEGO objects that fall from a height of about 30cm from a designated plane, and design a CNN-based recognition system that processes the impact sounds to determine the type of LEGO it belongs to. Recognizing that the fall of LEGO results in the main impact sound (i.e., only the sound at the moment of impact) and several subsequent sounds, we examine whether considering only the first impact sound or all sounds brings about better classification accuracies. We propose a compact two-dimensional CNN model, namely LegoNet, which is designed with a frame-level attention module at the input spectrogram and time-distributed fully-connected layers. Our experiments show that free-fall impact sounds can be used efficiently for accurate object recognition, and the proposed LegoNet, with a much smaller size, achieves better accuracy and robustness compared to baseline models. Also, using the whole sequence of impact sounds is more informative for LEGO classification than only considering the first impact sound. Moreover, it is found that utilizing data of specific object postures can help to improve the classifier’s performance in the case of small training data. The proposed approach can be employed as an extra module to build intelligent agents or object classification systems that require a rich understanding of the surrounding physical world.

Acoustic-based LEGO recognition using attention-based convolutional neural networks
Author:Van-Thuan Tran, Chia-Yang Wu, Wei-Ho Tsai
Year:2024
Source publication:Artificial Intelligence Review, Volume 57, article number 10, 2024
Subfield Highest percentage:99% Language and Linguistics #1 / 1126
https://www.sciencedirect.com/science/article/pii/S2352710224023970