We provide IT Staff Augmentation Services!

Big Data Solution Architect Resume

2.00/5 (Submit Your Rating)

SUMMARY:

  • I have 3+ years of experience in architecting and implementing big data and analytics systems using Cloudera and Confidential technologies with AWS and on premise.
  • I am a solution architect for Product domain and data science in Sainsbury’s DPP big data program. Previously, i was technical lead Confidential big data program (5 Mil $ budget), technical lead of Yapikredi bank DWH Modernization program and Next best action framework implementation. I worked in Ministry Of Education machine log data analytics on Confidential Big data appliance.
  • I have 360 degrees of experience in data management with big data, data architecture, data science, data warehouse design, data quality management, analytical datamart design, regulatory reporting, many CRM integration and reporting system design roles with stakeholder management in retail and financial industry for 12 years.

SKILLS:

File formats: XML, JSON, Avro, Parquet, sequence file, text file with Snappy, gzip, bzip compression algorithms.

Programming languages: Java, python, scala, R. Python Pandas, numpy, nltk

 Data mining: Confidential Advanced Analytics, R statistics, Jupyter, Weka

Data modeling: Relational modeling, Dimensional Kimball modeling, Erwin, Power designer, Sql developer studio. Master data management, Data governance

Data integration: Confidential ODI (10g/11g/12c) KM, Sdk development.

Operating systems: Unix, AIX, Linux, Windows XP/7/8, shell programming, Toad, sql developer

 Databases: Confidential, IBM DB2, SQL Server 2005,2008, AS/400, SYBASE, MySql

 Visualization & Reporting: OBIEE 3 - 4 years, Confidential Big data discovery 1.1, Endeca WAT, RTD.

PROFESSIONAL EXPERIENCE:

Confidential

Big data solution architect

Responsibilities:
  • Analyzing product data sources Product Enterprise services, external sources and BIW systems, designing staging, foundation and access layer data model.
  • Integration of analytics data and data science algorithm pipelines to Data lake.
  • Designing enterprise data models with Erwin, designing ingestion data flow and strategy, mapping sources to data model, parsing complex data types (XML, JSON) and designing data transformation pipelines.
  • Designing data ingestion batch and near real time data flows from (REST web service, JMS, SFTP) sources using Kafka and Spark to AWS S3 by applying obfuscation for PII security compliance requirements.
  • Designing mappings from source system to staging, information and access layers and designing data transformations pipelines with Spark and EDIE.
  • Processing XML and JSON documents with XPATH, lxml parsers.
  • Tuning spark data pipelines and optimizing Hive tables with partitioning, chunking, coalescing and compression strategies using Parquet/text formats with snappy compression.

Key technologies: Spark, Hadoop, cloudera, impala, Kerberos, java, python, hive, AWS s3, Erwin, parquet, snappy, XML, xpath, JSON, YAML, Oozie, Jira, confluence.

Confidential

Technical lead

Responsibilities:
  • I was Lead data architect of a 20+ million sterling project in Yapi Kredi Bank which is one of the top tier Banks in Turkey. I was leading the Data management and big data strategies working closely with data modelers, analysts, ETL team and stakeholders.
  • I delivered artifacts for standards and best practices on Confidential information reference architecture, logical physical data modeling and naming standards, data quality approach, conformed layer design, ETL best practices(type 1, SCD), reference data, referential integrity, physical implementation and storage tuning.
  • In this program, I also worked designing the solution for next best action framework for marketing using big data technologies. We proposed a solution with Confidential NOSQL and prepared a POC. Initial solution response time was 3-5 seconds, after tuning NOSQL data model and application we managed to get 0.05-o.25 second response time.

Key technologies: Confidential Nosql, Java, Sybase Power Designer, Logical modeling, Conceptual modeling, Kimball dimensional modeling, OEAF ( Confidential enterprise architecture framework).

Confidential

Technical lead

Responsibilities:
  • In this project I started to engage Confidential during Confidential ’s sales process. I performed big data discovery workshops, use-case generation/consolidation, architecture blueprint design
  • We designed Data Lake and implemented data ingestion, data transformation and visualization with Confidential Big data appliance stack with one colleague who has administration expertise.
  • Our use cases were; social Media (Twitter, Facebook, foursquare) ingestion with Kafka and Flume. Web clickstream data with Flume Kafka. Network security data ingestion with Flume, Confidential and Sql source ingestion with Sqoop, Oraoop and tuning. Web server log ingestion with Flume kafka, Analytical log ingestion with Kafka Flume, Spark data transformations. ATM log ingestion Flume, built Full-text search on Apache SOLR
  • I incorporated data ingestion and ETL patterns for Sqoop, Confidential loader for Hadoop, Confidential sql connector for Hadoop with Kerberos security to Orace Data integrator 12c tool.
  • We parsed and transformed log files with Spark, Hive Spark streaming and managed in Hive tables using Avro, JSON, Parquet, Sequence file and flat files with compressed formats Snappy, gzip and splittable block compression. For fast changing sources like social media we applied Avro schema evolution.
  • We used Confidential Big data Discovery 1.0 version for visualizations and Impala for ad-hoc querying.
  • For data science team we used R programming and Spark-Mlib to replace existing SAS enterprise miner models.
  • In Confidential ; Big data implementation, installing and configuring big data components (Flume, sqoop, Kafka, Hue workbooks).
  • We configured Sqoop, flume and kafka components, tuned Mapreduce, HDFS, Hive parameters, designed security with Kerberos, HDFS ACL, Sentry and enabled Confidential Big data connectors OLH, OSCH, XQUERY, ORAAH.
  • We built the Data lake architecture with HDFS folder, schema, file system management best practices.

Key technologies: Flume, Sqoop, Kafka, Hue, Mapreduce, HDFS, Hive, Impala, pyspark, HDFS ACL, Kerberos, Apache Sentry, Confidential big data connectors:OLH, OSCH, XQUERY, ORAAH, shell scripting.

Confidential

Technical lead

Responsibilities:
  • I was a developer in ministry of education’s big data implementation where we ingested data with flume and parsed tablet pc logs with hadoop and analyzed using hue and impala. Then integrated structured data to Confidential exadata reporting environment.

Confidential

DWH designer

Responsibilities:
  • In my role as DWH designer in Confidential, I was designing relational models, dimensional models, designing ETL with ODI, performance tuning and optimizing enterprise data warehouse.

Key Skills: Data warehouse, Confidential, Exadata, Dimensional modeling, Kimball, ODI 12c, XML, XSD, OBIEE 11g, Sql developer, Toad, ETL, Sql optimization, data migration, PL/SQL,JIRA, shell scripting

Confidential

Lead Migration Data Architect

Responsibilities:
  • I was lead architect of the Core Banking system development program in Confidential . We designed a new Banking system and migrated 3 legacy system to the new system. I reporting directly to head of IT

Key Skills: Data migration, Datapump, PL/SQL, external tables, Confidential transparent Gateway, Confidential, Exadata, Toad, ETL, Sql optimization, data migration, shell scripting 

We'd love your feedback!