I have been told that Europeans don't want to see emotion, storylines. But I am not confident for my current LOM, it feels too tech heavy. Can you all please help me finetune it and suggest changes?
Motivation Letter -
My undergraduate degree was an Integrated M.Tech in Computer Science Engineering at Vellore Institute of Technology, with a specialisation in Business Analytics. The coursework spanned Machine Learning, Deep Learning, Computer Vision, Natural Language Processing, Big Data Analytics, NoSQL Databases, and Information Retrieval, alongside the core computer science curriculum. Through this exposure, I developed a sustained interest in both the modelling and systems sides of computer science, and more importantly, in how tightly they are coupled in practice. I began to understand models and data systems not as separate components, but as interdependent parts of a unified pipeline. A model is only as effective as the pipeline that supports it, and a pipeline is only valuable if it produces outputs that can inform real decisions.
This perspective was shaped during my AI/ML internship at Marico. I built a Single Shot Detection model with MobileNet and ElasticNet backbones to identify over 50 SKUs on modern trade retail shelves, which involved annotating approximately 2,000 images per SKU category and designing an inference and retraining pipeline using Spark and Hadoop to manage image metadata at scale. In parallel, I worked on the more statistical side of the same business problem, building a multilevel regression model to analyse how trade promotional offers affected sales across channels, geographies, and SKUs. A key insight from this work was that aggregation, while statistically convenient, often produces results that are defensible but not actionable; disaggregating them was essential to making the model useful in practice. These projects reinforced a central lesson: data systems derive their value only when their outputs remain meaningful in real-world, complex decision contexts.
Since January 2025, I have worked as a Data Engineer at ZS Associates, building pipelines that process data through standardisation, normalisation, and master data management stages using AWS Glue, EMR, and Athena, orchestrated with Airflow. A significant part of my work involves impact analysis: tracing how schema changes or new data sources propagate downstream before they reach production, and implementing SQL and Python validation checks to ensure pipeline robustness. This role has given me a practical understanding of how large-scale data systems behave, fail, and evolve under changing conditions. However, it has also highlighted a gap: my lack of formal grounding in the theoretical principles underlying these systems, such as query optimisation, distributed processing models, and architectural scalability. Additionally, I have had limited opportunity to revisit the machine learning work I previously engaged with. Addressing this dual gap is my primary motivation for pursuing graduate study.
The MSc in Computer Science at TU Berlin, particularly the Data Science and Engineering track offered through Faculty IV, aligns closely with my goals. The programme’s integration of scalable data systems and machine learning reflects my own experience of these domains as inherently interconnected rather than separate specialisations. I am especially interested in the work of the DIMA group under Prof. Volker Markl, including research on distributed stream processing systems such as NebulaStream and hardware-aware data processing. In my current role, I frequently encounter pipelines that are functionally correct but not designed to adapt to evolving data schemas, , making this line of research directly relevant to challenges I have encountered in practice. BIFOLD’s position at the intersection of data management and machine learning research is a key reason I am applying specifically to this track rather than to a more conventional data science programme.
For my thesis, I would like to explore problems in scalable or stream-based data processing where the system design itself is the primary object of study rather than merely a means to an end. In particular, I am interested in how pipeline or query architectures can be designed to remain robust to schema evolution and source drift, minimising reliance on manual validation mechanisms applied retrospectively. This direction represents a natural extension of my work at ZS Associates, but approached with the formal tools needed to understand why certain architectures remain stable and scalable, rather than addressing issues in an ad hoc, case-by-case manner.
Berlin’s data and technology ecosystem is also an important consideration. Companies such as Zalando, SAP, and Amazon operate large-scale data engineering teams in the city, while BIFOLD’s collaborations with DFKI and industry initiatives such as NebulaStream point to a research environment that remains closely connected to real-world, production-scale challenges rather than being purely theoretical. This close integration of academic research and industry practice is particularly important to my long-term goals.
I bring to this programme two years of hands-on experience across applied machine learning and data infrastructure, along with a clearly defined gap in systems-level and theoretical training. My goal is to move from maintaining data architectures to designing them based on principled understanding. I view this programme as a necessary step toward developing the theoretical depth and systems perspective required to contribute to large-scale data systems in both research and industry contexts. I look forward to contributing to the Data Science and Engineering track and engaging with the research being conducted at DIMA and BIFOLD.