Repository navigation
Feat/pg vs neo4j benchmark - #3
Merged
Merged
Conversation
Neo4j 그래프(Product-CONTAINS->Ingredient-AFFECTS->Effect-RELATES_TO->Concern)와 동일한 개체/관계를 정규화된 RDB 스키마로 옮김. 다대다 관계는 연결 테이블로, 프로덕션 쿼리(neo4j_client.py)가 실제로 타는 조인/필터 컬럼 기준 인덱스 포함.
4EVR0-Server의 Postgres(5432, 세션 저장용)와 분리된 독립 인스턴스를 5433 포트로 띄워 벤치마크 데이터가 다른 용도와 섞이지 않게 함.
관계가 명확할수록 오히려 RDB가 유리할 수 있다는 논의, 스키마를 "그대로 가져간다"는 것의 의미, 연결 테이블이 그래프 흉내가 아니라 다대다 관계의 표준 RDB 표현이라는 결정 근거를 문서화.
기존 neo4j/data/ 무시 패턴과 동일하게, 벤치마크용 Postgres 컨테이너의 볼륨 마운트 경로(pg_experiment/data/)도 버전관리 대상에서 제외.
Neo4j는 멀티그래프라 같은 (ingredient, effect) 쌍에 관계가 여러 개 있을 수 있는데(affects.csv에 58쌍 중복 존재), (inci_name, effect_code)를 PK로 걸면 이 중복이 적재 시 거부되어 원본보다 적은 행을 갖게 되고 벤치마크가 Postgres에 유리하게 왜곡됨. BIGSERIAL PK로 바꿔 전체 행 유지.
csv/nodes, csv/edges의 neo4j import용 CSV를 그대로 읽어 동일한 원본 데이터를 Postgres에 적재. product 3122 / ingredient 3221 / effect 15 / concern 15 / contains 112966 / affects 5386 / relates_to 24 행 확인.
neo4j_client.py의 query_products_by_ingredients, query_ingredients_by_effects, query_path_by_effects와 동등한 결과를 내는 SQL 작성. Cypher 원본도 함께 보관해 로직이 갈라지지 않게 함.
검증 중 두 가지 실제 이슈를 발견/기록: 1. Postgres DB collation(en_US.utf8)이 한글 문자열을 Neo4j(코드포인트 기준)와 다르게 정렬해 top-N 결과 집합 자체가 달라짐 -> ORDER BY에 COLLATE "C" 적용해 해결. 2. query_path_by_effects는 원본 Cypher가 graph_score DESC 하나만으로 정렬해서 동점 구간에서 어느 행이 LIMIT 안에 들지가 원본부터 비결정적 -> 이 쿼리만 graph_score 분포로 비교(신원 비교 안 함).
프로덕션 쿼리 3개는 전부 1~2-hop이라, "hop이 늘어나면 그래프DB가 유리해지는 지점이 있는가"를 보려고 README의 전체 경로 (Product-CONTAINS->Ingredient-AFFECTS->Effect-RELATES_TO->Concern)를 쓰는 4-hop 쿼리를 실험용으로 추가. 실서비스는 안 씀(eval/RESULTS.md 참고).
동일 파라미터 시퀀스(seed 고정)를 양쪽 엔진에 그대로 실행해 워밍업 20회+ 측정 200회로 p50/p95/p99를 비교. 세션/커넥션은 재사용해 순수 쿼리 실행 시간만 측정(실 서비스도 드라이버가 커넥션을 풀링하는 것과 동일 조건).
정합성 검증(verify_parity.py) 원본 출력 - collation 버그 발견/수정 과정 포함 - 과 벤치마크(benchmark.py) 원본 출력을 그대로 첨부. 프로덕션 쿼리 3개 기준으로는 1-hop 둘은 Postgres가, 2-hop 하나는 Neo4j가 우세했고, hop-scaling 실험용 4-hop 쿼리에서는 Neo4j가 p99 기준 약 43배 앞섬 - hop이 늘어날수록 그래프DB가 유리해지는 패턴을 실측으로 확인.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RDB(Postgres) vs GraphDB(Neo4j) 성능 비교 실험
요약
4EVR0-Server/app/clients/neo4j_client.py)와 동일한 데이터(csv/nodes,csv/edges)를 Postgres에 동등한 스키마로 옮기고, 동등한 SQL로 포팅해 latency를 비교COLLATE "C"로 해결), 그리고 프로덕션 쿼리 자체의 기존 정렬 비결정성(path_by_effects)자세한 논의/과정은
pg_experiment/EXPERIMENT.md, 결과 원본 출력은pg_experiment/RESULTS.md참고.구성
pg_experiment/schema.sql— RDB 스키마 (다대다 관계는 연결 테이블로)pg_experiment/docker-compose.yml— 벤치마크 전용 Postgres 컨테이너(5433, 4EVR0-Server용과 분리)pg_experiment/load_csv.py— CSV → Postgres 적재pg_experiment/queries.py— Cypher 원본 + 포팅한 SQLpg_experiment/verify_parity.py— 두 엔진 결과 일치 검증pg_experiment/benchmark.py— latency 벤치마크 하네스pg_experiment/RESULTS.md,pg_experiment/results/latencies.json— 결과다루지 않은 것 (후속 과제)
SIMILAR_TO*1..3같은) — 지금 데이터엔 없어서 테스트 못 함, 그래프DB가 전형적으로 유리한 영역이라 후속 실험 후보Test plan
verify_parity.py로 3개 프로덕션 쿼리의 Cypher/SQL 결과 일치 확인benchmark.py로 4개 쿼리 x 200회 반복 latency 측정 및 결과 기록