Bigdata Engineer Resume
3.00/5 (Submit Your Rating)
FloridA
PROFESSIONAL SUMMARY
- Accomplished IT Professional with 7.5 years of IT experience in design and implementation of data warehousing, analytics, data integration and big data projects. This includes more than 4 years of experience in Big data Technologies.
- Strong communication and leadership skills with sound knowledge and practical experience in data warehousing concepts. A self - motivated hardworking team player with short learning curve and the constant zeal to learn more.
- Worked on multiple data science projects to build statistical regression model and model selection.
- Experience in predictive analytic procedures used in supervised learning (Regression, Neural Networks, Decision trees), unsupervised learning (Clustering-k-Means and Hierarchical, PCA).
- Experience in descriptive, exploratory, inferential, predictive modeling of the given dataset along with data modeling and recognizing key performance indicators (KPI).
- Good experience in predictive modeling, machine learning and data mining using python and R.
- Successfully worked and implemented multiple end-to-end Projects independently on various database applications depending on business partners requirements.
- Extensive Experience in data modeling, data warehouse and Ralph Kimball models with Star/Snowflake Schema Designs with analysis-definition, database design, testing, and implementation and Quality process.
- Successfully worked on leveraging multiple ETL technologies like datastage, hadoop and teradata to develop ETL applications to achieve maximum efficiency and optimum resource utilization.
- Consulted with business partners and made recommendations to improve the effectiveness of Big Data systems, descriptive analytics systems, and prescriptive analytics systems.
- Proficient in creating new data collection systems that optimize data management, capturing, delivery and quality
- Working knowledge of Big Data Analytics, Hadoop ecosystems (Hadoop, Hive).
- Experience in utilizing HIVE for working with data stored in the Hadoop file system (HDFS)
- Sound understanding of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, and Map Reduce concepts.
- Worked on developing big data solutions using Apache Hadoop ecosystem and tools like Hive, Spark, Sqoop, Pig, Kafka, Python, Oozie, NIFI.
- Designed and Implemented data streaming applications to produce and consume data feeds using Kafka.
- Successfully built and migrated multiple legacy application into hadoop environment with added efficiency and performance.
- Extensive experience in ETL/ELT methodologies supporting Data Migration, Data Transformation and Data Cleansing of structured, unstructured and semi structured data formats using ETL tools (Informatica Power Center, DataStage).
TECHNICAL SKILLS:
BigData Tool: Hive, Pig, Oozie, Sqoop, Impala, SAS, R, Kafka
Programing: SQL, Unix Shell Scripting, Phyton
ETL Tools: IBM DataStage, Informatica
DataBases: Teradata 13.0, 14.1, MySQL, NoSQL MongoDB
Schedulers: Autosys, Oozie
Versioning: GITHUB, SVN tortoise
PROFESSIONAL EXPERIENCE:
BigData Engineer
Confidential, Florida
Responsibilities:
- Perform End-to-end development activities, from requirement gathering and analysis, to system design, coding and testing.
- Systematically compiled requirements and performed in-depth impact analysis
- Develop various Teradata utilities like Mload, Bteq, FastExport etc. for required data transfer of various applications.
- Query Tuning and Index optimization for various complex SQL queries in production.
- Extract data to HDFS from Teradata/Oracle using Sqoop(1.4) for customer journey index application which feeds data to customer churn model.
- Import/Export data with HDFS from Amazon Redshift(AWS) using Sqoop
- Load data files from UNIX server to HDFS for loading into HIVE database.
- Transform and analyze the data using HIVE (0.14) & PIG (0.12).
- Perform performance tuning on HIVE by using concepts like partitioning and merge multiple small files etc.
- Ingest (import/export) data to and from hdfs into rdbms using Sqoop for different kinds of file formats
- Develop & Schedule Oozie(4.1) workflows for processing data.
- Real time data streaming using Apache Kafka and NIFI
Data Engineer
Confidential, Jacksonville, Florida
Responsibilities:
- Gathered requirements from client partners for application Development and implement end to end project with zero defects.
- Implemented slowly changing Dimension logics in the mapping to effectively handle change data capture which is typical in data warehousing systems.
- Developed automated process to perform data quality checks between two systems and generate quality reports via data stage and Hadoop technologies.
- Leveraged multiple ETL technologies like Datastage, Hadoop and Teradata to develop ETL applications to achieve maximum efficiency and optimum resource utilization.
- Build Data Model by analyzing the table structures to reduce Data Redundancy.
- Involved in Several POC on Cloudera Hadoop converting small, medium, complex legacy functionality into Hadoop
- Connected to hive backend Postgres sql and extracted critical hive metadata information and populated system dictionary tables in hive.
- Develop & Schedule Oozie(4.1) workflows for processing data.
- Load data files from UNIX server to HDFS for loading into HIVE database.
Teradata Applications ETL Developer
Confidential
Responsibilities:
- Managed around 450 applications in Teradata Data Warehouse (Called The W) across different LOBs under ECIO domain in Global Technologies and operations.
- Planning and execution of teradata hardware/Software upgrades, Disaster Recovery exercises and technology Migrations across the years and coordinated the recovery for enterprise Applications post upgrade to bring the system back to BAU.
- Worked on application enhancements running in production environment to improve efficiency involving better run duration and resource utilization.
- Performed ETL operations on raw data from various sources namely mainframes, Datastage and Hadoop.
- Address user tickets and queries on the data and the applications with the business logic behind them.
- Worked as support level 2 & 3 analyst in Datastage and Informatica platform applications. Provided 24/7 support which involves responsibilities of resolving the abends under predefined SLA’s.
- Did performance analysis on both Datastage and target Teradata systems by analyzing for various bottlenecks and implementing indexes, collect stats and query tuning.
- Worked on various value add automations resulting in saving both in terms of man-hours as well as the cost savings.
- Query Tuning and Index optimization for various complex SQL queries in production.
