Bigdata Developer Resume
5.00/5 (Submit Your Rating)
CA
SUMMARY:
- Over 9+ years of experience in building distributed, scalable and complex applications using Hadoop Eco System, RDBMS, Spark and Informatica
- Around 4 years of experience in Hadoop ecosystems: Hive, Sqoop, Hbase and Flume
- 2+ years of experience in Spark using pyspark and scala
- Expertise in Map - Reduce frameworks with structured & unstructured data
- Well Versed in Apache Spark Framework - Data Frames, SQL, JSON, PARQUET files
- Good knowledge of Hadoop Architecture and various components such as HDFS, Map Reduce, Hive, Job Tracker, Task Tracker, Name Node, Data Node and YARN
- Expertise in SparkQL/HiveQL and implemented with best practices.
- Worked on exporting/importing data using Sqoop from HDFS to Relational Database Systems and vice-versa
- Proficiency in distributed query-processing tools like Hive, Spark SQL
- Used tools like SQOOP, Kafka to ingest data into Hadoop.
- Good understanding of cloud configuration in Amazon web services AWS
- Excellent understanding of Tableau visualization and experience in building dashboards
- Working knowledge in connecting Power Center to Teradata using FASTLOAD, BULKLOAD, MULTILOAD and TPUMP
- Involved in full life cycle, which includes Data Analysis, Data Validation, Data Verification, Data Cleansing, Data Completeness and identifying data mismatch., and execution of the Process and Documentation.
- Well versed with the concepts, design and development of DataMart s, Data Warehouse using Star Schema and Snowflake Schema and also implementing Decision Support System.
- Excellent knowledge in database design, relational integrity constraints, OLAP, OLTP, Cubes and normalizations.
- Extensive experience in DW Tools like Informatica, Oracle, TERADATA
- Strong problem solving skills, quick learner and able to work independently as well as a team member of varying sized teams.
- Ability to plan, manage, motivate and work efficiently as an independent or collaboratively in a team.
- Self-motivated, enthusiastic and always keen to learn new methodologies and techniques
TECHNICAL SKILLS:
Big Data / Hadoop: HDFS, Map Reduce, Hive, Spark, Hbase, and Sqoop.
ETL Tools: Informatica 9.5.1 and Informatica Data Quality 8.6.1 , Teradata utilities
Reporting Tools: Tableau, Business Objects XI
Languages: PL/SQL, Python, Unix shell scripting
Database and Tools: Oracle 11g, Teradata V14, SQL Server 2007Oracle SQL developer 2.1
Operating Systems: Linux,UNIX and Windows
Other Tools: Eclipse, SBT,GIT
PROFESSIONAL EXPERIENCE:
BigData Developer
Confidential, CA
Responsibilities:
- Built data pipelines to Load and transform large sets of structured, semi structured and unstructured data.
- Imported data from HDFS into Hive using HiveQL
- Involved in creating Hive tables, loading and analyzing data using hive queries
- Created Hive Partitioned and Bucketed tables to improve performance.
- Developed a SQOOP Import Job, Shell Script & CRONJOB for importing data into HDFS
- Used Tableau for visualization and building dashboards
- To improve performance and optimization of the existing algorithms, explored different components like Spark Context, Spark-SQL, Data Frame, Pair RDD's, accumulators
- Processed millions of records using Hadoop jobs
- Implemented Spark code using Python for RDD transformations & actions in Spark application
- Built reusable Hive UDF libraries for business requirements
- Working with the leadership to understand scope, derive estimates, schedule, allocate work, manage tasks/projects, present status updates to IT and business leaders as required.
- Define and contribute to development of standards, guidelines, design patterns and common development frameworks & components.
Data engineer
Confidential, CA
Responsibilities:
- Translate ETL requirements into formalized designs and mapping documents.
- Handled importing data from various data sources using Sqoop. Built pipelines to transform data using MapReduce/spark to load data into HDFS.
- To improve performance and optimization of the existing algorithms, explored different components like Spark-SQL, Data Frame, Pair RDD's, accumulators
- Implemented Spark code using Scala for RDD transformations & actions in Spark application
- Hands-on experience with systems-building languages such as MapReduce in Python
- Implemented hive tables on different file format like ORC and Parquet for efficient retrieval of needed columns
- Optimized queries to use different types of joins like map side and bucket joins for efficient query retrievals
- Involved in running Hadoop jobs for processing millions of records and compression techniques
- Developed multiple MapReduce jobs for data cleaning and pre-processing
- Integrate external APIs (Salesforce, Google APIs, etc.) and ETL into data marts using Python
- Built aggregates for canned reports using Hive
- Implemented Broadcast & Accumulators variables in the application.
- Developed Fast export scripts to spool files for third parties
- Involvement in implementation of BTEQ and Bulk load jobs.
- Built ETL logic using Teradata bteq for marketing data models to measure campaign effectiveness
- Created Different Fastload, Multiload and BTEQ Scripts to facilitate the ETL Processes
- TPT Execution is handled in Wrapper scripts and few occasions’s using Informatica loader’s to load data to different data bases for reporting purposes.
- Expertise in performance tuning the user queries. execution of frequently used SQL operations and improve the performance.
- Perform unit test cases and run validation on the views to ensure data quality.
- Created complex mappings using Sql and Flagged the record using update strategy for populating the desired slowly changing dimension tables.
- Involved in discussions with Business Analysts and Business Users to design technical specification documents for Data warehouse
- Followed Agile Methodology especially SCRUM software development process throughout development, QA and UAT
- Leadership to advocate Dev team with User and other Technical Team Members.
- Performed sizing for various EDW activities
- Mentoring the developers and aiding the business and testing teams with the testing process.
- Audit architectural design and framework aka validation rule engine
- To find data anomalies implemented scoring based audit.
- To find data holes, created alerts if the data volumes are not in the desired threshold limits
- Extensively researched and fixed error events /quality issues in production/QA environment.
DWH ETL Consultant
Confidential, CA
Responsibilities:
- Worked in the full life cycle development of the warehouse from analysis through implementation which includes data modeling, Extract, Transform and Load process and Report development
- Envisioned data model for future needs by keeping in mind the current needs and growth
- Built Data quality screens in Informatica developer and designer to automate data cleansing in Data Warehouse
- Followed Agile Methodology especially SCRUM software development process throughout development, QA and UAT
- Strong business knowledge in Transportation domain especially Container Leasing business which involves Daily and Monthly Fleet Status, Daily Activity, Accounting - General Ledger, Revenue and Expense and Product Management
- Responsible for Extraction, Transformation & Loading data in Tables using Teradata tools like bteq, multiload, fastload and Informatica Power Center 9.1/8.6 and Oracle PL/SQL.
- Developed multiple Workflows for initial / historical data loading, using sessions, worklets, commands, decisions, events and timers
- Designed and developed the ETL framework, including coordinating/executing.
- Helped QA team to perform parallel testing with the existing system.
- Worked closely with testers for Regression testing between cycles and releases.
- Coordinating the “certification” and sign-off of the release.
- Mentoring the developers and aiding the business and testing teams with the testing process.
- Used EXPLAIN PLAN to tune queries for better performance and also Extensive Usage of Indexes
- Worked extensively on Teradata SQL query tuning using sliding window mechanism, temporary tables, collect stats, skew indexes and join indexes
- Lead developer in 3NF database modeling and database optimization
- Trained users on Business Objects especially in creating reports, variables, import external data sources, export to cms, access. etc.
